paper-with-me

Papers

Spanish and LLM Benchmarks: is MMLU Lost in Translation?

2024-05-28 · Irene Plaza, Nina Melero, Cristina del Pozo, Javier Conde, Pedro Reviriego, Marina Mayor-Rocher, María Grandury

The evaluation of Large Language Models (LLMs) is a key element in their continuous improvement process and many benchmarks have been developed to assess the performance of LLMs in different tasks and topics. As LLMs become adopted worldwide, evaluating them in languages other than English is increasingly important. However, most LLM benchmarks are simply translated using an automated tool and then run in the target language. This means that the results depend not only on the LLM performance in that language but also on the quality of the translation. In this paper, we consider the case of the well-known Massive Multitask Language Understanding (MMLU) benchmark. Selected categories of the benchmark are translated into Spanish using Azure Translator and ChatGPT4 and run on ChatGPT4. Next, the results are processed to identify the test items that produce different answers in Spanish and English. Those are then analyzed manually to understand if the automatic translation caused the change. The results show that a significant fraction of the failing items can be attributed to mistakes in the translation of the benchmark. These results make a strong case for improving benchmarks in languages other than English by at least revising the translations of the items and preferably by adapting the tests to the target language by experts.

📄 PDF Abstract BibTeX arXiv:2406.17789

Code (0)

등록된 구현이 없습니다.

Tasks

MMLUTranslation

Similar Papers 제목 키워드 기반

Lost in Translation: Analysis of Information Loss During Machine Translation Between Polysynthetic and Fusional Languages

2018-07-01 · COLING 2018 8 · Manuel Mager, Elisabeth Mager, Alfonso Medina-Urrea, Ivan Meza 외

Machine translation from polysynthetic to fusional languages is a challenging task, which gets further complicated by the limited amount of parallel text available. Thus, translation performance is far from the state of …

Machine TranslationTranslation

Multi-lingual Functional Evaluation for Large Language Models

2025-06-25 · Victor Ojewale, Inioluwa Deborah Raji, Suresh Venkatasubramanian

Multi-lingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM. However, these evaluations often fail to provide an adequate understanding of the practi…

BelebeleInstruction FollowingMathMMLU

None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks

2025-02-18 · Eva Sánchez Salido, Julio Gonzalo, Guillermo Marco

In LLM evaluations, reasoning is often distinguished from recall/memorization by performing numerical variations to math-oriented questions. Here we introduce a general variation method for multiple-choice questions that…

MathMemorizationMMLUMultiple-choice

Neural Machine Translation of Text from Non-Native Speakers

2018-08-19 · NAACL 2019 6 · Antonios Anastasopoulos, Alison Lui, Toan Nguyen, David Chiang

Neural Machine Translation (NMT) systems are known to degrade when confronted with noisy data, especially when the system is trained only on clean data. In this paper, we show that augmenting training data with sentences…

Machine TranslationNMTTranslation

Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts

2024-03-17 · Michael Saxon, Yiran Luo, Sharon Levy, Chitta Baral 외

Benchmarks of the multilingual capabilities of text-to-image (T2I) models compare generated images prompted in a test language to an expected image distribution over a concept set. One such benchmark, "Conceptual Coverag…

Translation