paper-with-me

홈 › Papers

Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ

2024-03-06 · Carolin Holtermann, Paul Röttger, Timm Dill, Anne Lauscher

Large language models (LLMs) need to serve everyone, including a global majority of non-English speakers. However, most LLMs today, and open LLMs in particular, are often intended for use in just English (e.g. Llama2, Mistral) or a small handful of high-resource languages (e.g. Mixtral, Qwen). Recent research shows that, despite limits in their intended use, people prompt LLMs in many different languages. Therefore, in this paper, we investigate the basic multilingual capabilities of state-of-the-art open LLMs beyond their intended use. For this purpose, we introduce MultiQ, a new silver standard benchmark for basic open-ended question answering with 27.4k test questions across a typologically diverse set of 137 languages. With MultiQ, we evaluate language fidelity, i.e. whether models respond in the prompted language, and question answering accuracy. All LLMs we test respond faithfully and/or accurately for at least some languages beyond their intended use. Most models are more accurate when they respond faithfully. However, differences across models are large, and there is a long tail of languages where models are neither accurate nor faithful. We explore differences in tokenization as a potential explanation for our findings, identifying possible correlations that warrant further investigation.

📄 PDF Abstract BibTeX arXiv:2403.03814

Code (1)

paul-rottger/multiq 공식 구현

Tasks

Open-Ended Question AnsweringQuestion Answering

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

M2G-Eval: Enhancing and Evaluating Multi-granularity Multilingual Code Generation

2025-12-27 · Fanglin Xu, Wei Zhang, Jian Yang, Guo Chen 외 arxiv

The rapid advancement of code large language models (LLMs) has sparked significant research interest in systematically evaluating their code generation capabilities, yet existing benchmarks predominantly assess models at…

Code Generation

MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language

2025-05-20 · Seyoung Song, Seogyeong Jeong, Eunsu Kim, Jiho Jin 외

Evaluating text generation capabilities of large language models (LLMs) is challenging, particularly for low-resource languages where methods for direct assessment are scarce. We propose MUG-Eval, a novel framework that …

Text Generation

M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models

2024-05-24 · Hongyu Wang, Jiayu Xu, Senwei Xie, Ruiping Wang 외

Multilingual multimodal reasoning is a core component in achieving human-level intelligence. However, most existing benchmarks for multilingual multimodal reasoning struggle to differentiate between models of varying per…

Multimodal Reasoning

MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

2025-03-13 · Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng 외

Traditional benchmarks struggle to evaluate increasingly sophisticated language models in multilingual and culturally diverse contexts. To address this gap, we introduce MMLU-ProX, a comprehensive multilingual benchmark …

Language Model EvaluationLanguage ModelingLanguage ModellingLarge Language Model+1

Towards Multilingual LLM Evaluation for European Languages

2024-10-11 · Klaudia Thellmann, Bernhard Stadler, Michael Fromm, Jasper Schulze Buschhoff 외

The rise of Large Language Models (LLMs) has revolutionized natural language processing across numerous languages and tasks. However, evaluating LLM performance in a consistent and meaningful way across multiple European…

ARCGSM8KHellaSwagMMLU+1