paper-with-me

홈 › Papers

AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese

2026-03-27 · Afonso Simplício, Gonçalo Vinagre, Miguel Moura Ramos, Diogo Tavares, Rafael Ferreira, Giuseppe Attanasio, Duarte M. Alves, Inês Calvo, Inês Vieira, Rui Guerra, James Furtado, Beatriz Canaverde, Iago Paulo, Vasco Ramos, Diogo Glória-Silva, Miguel Faria, Marcos Treviso, Daniel Gomes, Pedro Gomes, David Semedo, André Martins, João Magalhães arxiv

Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant's linguistic and cultural nuances. We introduce AMALIA, a fully open LLM that prioritizes pt-PT by using more high-quality pt-PT data during both the mid- and post-training stages. To evaluate pt-PT more faithfully, we release a suite of pt-PT benchmarks that includes translated standard tasks and four new datasets targeting pt-PT generation, linguistic competence, and pt-PT/pt-BR bias. Experiments show that AMALIA matches strong baselines on translated benchmarks while substantially improving performance on pt-PT-specific evaluations, supporting the case for targeted training and native benchmarking for European Portuguese.

📄 PDF Abstract BibTeX arXiv:2603.26511

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

2026-06-17 · Diogo Glória-Silva, João Cardeira, Manuel Letras da Luz, Afonso Simplício 외 arxiv

Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open-source multimodal models, which either conflate it with Brazilian Portug…

Validity of LLMs as data annotators: AMALIA on authority

2026-07-09 · Manuel Pita arxiv

A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competiti…

GPT4All: An Ecosystem of Open Source Compressed Language Models

2023-11-06 · Yuvanesh Anand, Zach Nussbaum, Adam Treat, Aaron Miller 외

Large language models (LLMs) have recently achieved human-level performance on a range of professional and academic benchmarks. The accessibility of these models has lagged behind their performance. State-of-the-art LLMs…

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

2025-03-18 · Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 외

Inference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are c…

reinforcement-learningReinforcement Learning

Nomic Embed: Training a Reproducible Long Context Text Embedder

2024-02-02 · Zach Nussbaum, John X. Morris, Brandon Duderstadt, Andriy Mulyar

This technical report describes the training of nomic-embed-text-v1, the first fully reproducible, open-source, open-weights, open-data, 8192 context length English text embedding model that outperforms both OpenAI Ada-0…