paper-with-me

Papers

Spanish Pre-trained BERT Model and Evaluation Data

2023-08-06 · José Cañete, Gabriel Chaperon, Rodrigo Fuentes, Jou-Hui Ho, Hojin Kang, Jorge Pérez

The Spanish language is one of the top 5 spoken languages in the world. Nevertheless, finding resources to train or evaluate Spanish language models is not an easy task. In this paper we help bridge this gap by presenting a BERT-based language model pre-trained exclusively on Spanish data. As a second contribution, we also compiled several tasks specifically for the Spanish language in a single repository much in the spirit of the GLUE benchmark. By fine-tuning our pre-trained Spanish model, we obtain better results compared to other BERT-based models pre-trained on multilingual corpora for most of the tasks, even achieving a new state-of-the-art on some of them. We have publicly released our model, the pre-training data, and the compilation of the Spanish benchmarks.

📄 PDF Abstract BibTeX arXiv:2308.02976

Code (2)

dccuchile/beto 공식 구현 tf
josecannete/spanish-corpora 공식 구현

Tasks

Language ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

RoBERTuito: a pre-trained language model for social media text in Spanish

2021-11-18 · LREC 2022 6 · Juan Manuel Pérez, Damián A. Furman, Laura Alonso Alemany, Franco Luque

Since BERT appeared, Transformer language models and transfer learning have become state-of-the-art for Natural Language Understanding tasks. Recently, some works geared towards pre-training specially-crafted models for …

Language ModelingLanguage ModellingNatural Language UnderstandingTransfer Learning

MarIA: Spanish Language Models

2021-07-15 · Asier Gutiérrez-Fandiño, Jordi Armengol-Estapé, Marc Pàmies, Joan Llop-Palao 외

This work presents MarIA, a family of Spanish language models and associated resources made available to the industry and the research community. Currently, MarIA includes RoBERTa-base, RoBERTa-large, GPT2 and GPT2-large…

Extractive Question-AnsweringQuestion Answering

Sequence-to-Sequence Spanish Pre-trained Language Models

2023-09-20 · Vladimir Araujo, Maria Mihaela Trusca, Rodrigo Tufiño, Marie-Francine Moens

In recent years, significant advancements in pre-trained language models have driven the creation of numerous non-English language variants, with a particular emphasis on encoder-only and decoder-only architectures. Whil…

DecoderGenerative Question AnsweringNatural Language UnderstandingQuestion Answering+1

MEL: Legal Spanish Language Model

2025-01-27 · David Betancur Sánchez, Nuria Aldama García, Álvaro Barbero Jiménez, Marta Guerrero Nieto 외

Legal texts, characterized by complex and specialized terminology, present a significant challenge for Language Models. Adding an underrepresented language, such as Spanish, to the mix makes it even more challenging. Whi…

Language ModelingLanguage Modellingmodel

Evaluation Benchmarks for Spanish Sentence Representations

2022-04-15 · LREC 2022 6 · Vladimir Araujo, Andrés Carvallo, Souvik Kundu, José Cañete 외

Due to the success of pre-trained language models, versions of languages other than English have been released in recent years. This fact implies the need for resources to evaluate these models. In the case of Spanish, t…

Language ModelingLanguage ModellingSentence