paper-with-me

Papers

Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate

2025-05-28 · Ashim Gupta, Maitrey Mehta, Zhichao Xu, Vivek Srikumar

Large language models (LLMs) provide detailed and impressive responses to queries in English. However, are they really consistent at responding to the same query in other languages? The popular way of evaluating for multilingual performance of LLMs requires expensive-to-collect annotated datasets. Further, evaluating for tasks like open-ended generation, where multiple correct answers may exist, is nontrivial. Instead, we propose to evaluate the predictability of model response across different languages. In this work, we propose a framework to evaluate LLM's cross-lingual consistency based on a simple Translate then Evaluate strategy. We instantiate this evaluation framework along two dimensions of consistency: information and empathy. Our results reveal pronounced inconsistencies in popular LLM responses across thirty languages, with severe performance deficits in certain language families and scripts, underscoring critical weaknesses in their multilingual capabilities. These findings necessitate cross-lingual evaluations that are consistent along multiple dimensions. We invite practitioners to use our framework for future multilingual LLM benchmarking.

📄 PDF Abstract BibTeX arXiv:2505.21999

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

KoBE: Knowledge-Based Machine Translation Evaluation

2020-09-23 · Findings of the Association for Computational Linguistics 2020 · Zorik Gekhman, Roee Aharoni, Genady Beryozkin, Markus Freitag 외

We propose a simple and effective method for machine translation evaluation which does not require reference translations. Our approach is based on (1) grounding the entity mentions found in each source sentence and cand…

Machine TranslationSentenceTranslation

Beyond English: Evaluating Automated Measurement of Moral Foundations in Non-English Discourse with a Chinese Case Study

2025-02-04 · Calvin Yixiang Cheng, Scott A Hale

This study explores computational approaches for measuring moral foundations (MFs) in non-English corpora. Since most resources are developed primarily for English, cross-linguistic applications of moral foundation theor…

Machine TranslationTransfer Learning

Learning Multilingual Sentence Representations with Cross-lingual Consistency Regularization

2023-06-12 · Pengzhi Gao, Liwen Zhang, Zhongjun He, Hua Wu 외

Multilingual sentence representations are the foundation for similarity-based bitext mining, which is crucial for scaling multilingual neural machine translation (NMT) system to more languages. In this paper, we introduc…

DecoderMachine TranslationNMTSentence+1

NICT‘s Submission To WAT 2020: How Effective Are Simple Many-To-Many Neural Machine Translation Models?

2020-12-01 · AACL (WAT) 2020 12 · Raj Dabre, Abhisek Chakrabarty

In this paper we describe our team‘s (NICT-5) Neural Machine Translation (NMT) models whose translations were submitted to shared tasks of the 7th Workshop on Asian Translation. We participated in the Indic language mult…

Machine TranslationNMTTranslation

Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages

2026-03-18 · Yue Zhao, Jiatao Gu, Paloma Jeretič, Weijie Su arxiv

Understanding the distance between human languages is central to linguistics, anthropology, and tracing human evolutionary history. Yet, while linguistics has long provided rich qualitative accounts of cross-linguistic v…

Machine Translation