paper-with-me

홈 › Papers

Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions

2026-02-05 · Léo Labat, Etienne Ollion, François Yvon arxiv

Multiple-Choice Questions (MCQs) are often used to assess knowledge, reasoning abilities, and even values encoded in large language models (LLMs). While the effect of multilingualism has been studied on LLM factual recall, this paper seeks to investigate the less explored question of language-induced variation in value-laden MCQ responses. Are multilingual LLMs consistent in their responses across languages, i.e. behave like theoretical polyglots, or do they answer value-laden MCQs depending on the language of the question, like a multitude of monolingual models expressing different values through a single model? We release a new corpus, the Multilingual European Value Survey (MEVS), which, unlike prior work relying on machine translation or ad hoc prompts, solely comprises human-translated survey questions aligned in 8 European languages. We administer a subset of those questions to over thirty multilingual LLMs of various sizes, manufacturers and alignment-fine-tuning status under comprehensive, controlled prompt variations including answer order, symbol type, and tail character. Our results show that while larger, instruction-tuned models display higher overall consistency, the robustness of their responses varies greatly across questions, with certain MCQs eliciting total agreement within and across models while others leave LLM answers split. Language-specific behavior seems to arise in all consistent, instruction-fine-tuned models, but only on certain questions, warranting a further study of the selective effect of preference fine-tuning.

📄 PDF Abstract BibTeX arXiv:2602.05932

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Are Large Language Models Consistent over Value-laden Questions?

2024-07-03 · Jared Moore, Tanvi Deshpande, Diyi Yang

Large language models (LLMs) appear to bias their survey answers toward certain values. Nonetheless, some argue that LLMs are too inconsistent to simulate particular values. Are they? To answer, we first define value con…

Multiple-choice

Code-Mixing on Sesame Street: Dawn of the Adversarial Polyglots

2021-03-17 · NAACL (CALCS) 2021 6 · Samson Tan, Shafiq Joty

Multilingual models have demonstrated impressive cross-lingual transfer performance. However, test sets like XNLI are monolingual at the example level. In multilingual communities, it is common for polyglots to code-mix …

Cross-Lingual TransferXLM-R

Polyglot or Not? Measuring Multilingual Encyclopedic Knowledge in Foundation Models

2023-05-23 · Tim Schott, Daniel Furman, Shreshta Bhat

In this work, we assess the ability of foundation models to recall encyclopedic knowledge across a wide range of linguistic contexts. To support this, we: 1) produce a 20-language dataset that contains 303k factual assoc…

counterfactualRetrieval

The Thieves on Sesame Street are Polyglots - Extracting Multilingual Models from Monolingual APIs

2020-11-01 · EMNLP 2020 11 · Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, Richard Socher

Pre-training in natural language processing makes it easier for an adversary with only query access to a victim model to reconstruct a local copy of the victim by training with gibberish input data paired with the victim…

CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models

2024-05-22 · Giada Pistilli, Alina Leidinger, Yacine Jernite, Atoosa Kasirzadeh 외

This paper introduces the "CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal impacts" dataset, designed to evaluate the social and cultural variation of Large Language Models (LLMs) across multiple langu…