paper-with-me

Papers

EuroBERT: Scaling Multilingual Encoders for European Languages

2025-03-07 · Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Duarte M. Alves, André Martins, Ayoub Hammal, Caio Corro, Céline Hudelot, Emmanuel Malherbe, Etienne Malaboeuf, Fanny Jourdan, Gabriel Hautreux, João Alves, Kevin El-Haddad, Manuel Faysse, Maxime Peyrard, Nuno M. Guerreiro, Patrick Fernandes, Ricardo Rei, Pierre Colombo

General-purpose multilingual vector representations, used in retrieval, regression and classification, are traditionally obtained from bidirectional encoder models. Despite their wide applicability, encoders have been recently overshadowed by advances in generative decoder-only models. However, many innovations driving this progress are not inherently tied to decoders. In this paper, we revisit the development of multilingual encoders through the lens of these advances, and introduce EuroBERT, a family of multilingual encoders covering European and widely spoken global languages. Our models outperform existing alternatives across a diverse range of tasks, spanning multilingual capabilities, mathematics, and coding, and natively supporting sequences of up to 8,192 tokens. We also examine the design decisions behind EuroBERT, offering insights into our dataset composition and training pipeline. We publicly release the EuroBERT models, including intermediate training checkpoints, together with our training framework.

📄 PDF Abstract BibTeX arXiv:2503.05500

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

XplaiNLP at CheckThat! 2025: Multilingual Subjectivity Detection with Finetuned Transformers and Prompt-Based Inference with Large Language Models

2025-09-15 · Ariana Sahitaj, Jiaao Li, Pia Wenzel Neves, Fedor Splitt 외 arxiv

This notebook reports the XplaiNLP submission to the CheckThat! 2025 shared task on multilingual subjectivity detection. We evaluate two approaches: (1) supervised fine-tuning of transformer encoders, EuroBERT, XLM-RoBER…

SindBERT, the Sailor: Charting the Seas of Turkish NLP

2025-10-24 · Raphael Schmitt, Stefan Schweter arxiv

Transformer models have revolutionized NLP, yet many morphologically rich languages remain underrepresented in large-scale pre-training efforts. With SindBERT, we set out to chart the seas of Turkish NLP, providing the f…

Linguistic AcceptabilityPart-Of-Speech Tagging

EuroLLM: Multilingual Language Models for Europe

2024-09-24 · Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro 외

The quality of open-weight LLMs has seen significant improvement, yet they remain predominantly focused on English. In this paper, we introduce the EuroLLM project, aimed at developing a suite of open-weight multilingual…

Machine Translation

Towards Multilingual LLM Evaluation for European Languages

2024-10-11 · Klaudia Thellmann, Bernhard Stadler, Michael Fromm, Jasper Schulze Buschhoff 외

The rise of Large Language Models (LLMs) has revolutionized natural language processing across numerous languages and tasks. However, evaluating LLM performance in a consistent and meaningful way across multiple European…

ARCGSM8KHellaSwagMMLU+1

Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs

2024-09-30 · Mehdi Ali, Michael Fromm, Klaudia Thellmann, Jan Ebert 외

We present two multilingual LLMs designed to embrace Europe's linguistic diversity by supporting all 24 official languages of the European Union. Trained on a dataset comprising around 60% non-English data and utilizing …

ARCDiversityHellaSwagMMLU+1