paper-with-me

Papers

Quantifying Language Disparities in Multilingual Large Language Models

2025-08-23 · Songbo Hu, Ivan Vulić, Anna Korhonen arxiv

Results reported in large-scale multilingual evaluations are often fragmented and confounded by factors such as target languages, differences in experimental setups, and model choices. We propose a framework that disentangles these confounding variables and introduces three interpretable metrics--the performance realisation ratio, its coefficient of variation, and language potential--enabling a finer-grained and more insightful quantification of actual performance disparities across both (i) models and (ii) languages. Through a case study of 13 model variants on 11 multilingual datasets, we demonstrate that our framework provides a more reliable measurement of model performance and language disparities, particularly for low-resource languages, which have so far proven challenging to evaluate. Importantly, our results reveal that higher overall model performance does not necessarily imply greater fairness across languages.

📄 PDF Abstract BibTeX arXiv:2508.17162

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Systematic Study of Performance Disparities in Multilingual Task-Oriented Dialogue Systems

2023-10-19 · Songbo Hu, Han Zhou, Moy Yuan, Milan Gritta 외

Achieving robust language technologies that can perform well across the world's many languages is a central goal of multilingual NLP. In this work, we take stock of and empirically analyse task performance disparities th…

Language ModelingLanguage ModellingMultilingual NLPTask-Oriented Dialogue Systems

Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation

2024-04-12 · Haozhe Zhao, Zefan Cai, Shuzheng Si, Liang Chen 외

Large-scale multilingual Pretrained Language Models (mPLMs) yield impressive performance on cross-language tasks, yet significant performance disparities exist across different languages within the same mPLM. Previous st…

Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources

2022-11-28 · Xinyan Velocity Yu, Akari Asai, Trina Chatterjee, Junjie Hu 외

While the NLP community is generally aware of resource disparities among languages, we lack research that quantifies the extent and types of such disparity. Prior surveys estimating the availability of resources based on…

Zoom In Disparities in Healthcare LLM Q&A

2025-10-20 · Ipek Baris Schlicht, Burcu Sayin, Zhixue Zhao, Frederik M. Labonté 외 arxiv

Equitable access to reliable health information is vital when integrating AI into healthcare. Yet, information quality varies across languages, raising concerns about the reliability and consistency of multilingual Large…

M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks

2024-07-04 · Florian Schneider, Sunayana Sitaram

Since the release of ChatGPT, the field of Natural Language Processing has experienced rapid advancements, particularly in Large Language Models (LLMs) and their multimodal counterparts, Large Multimodal Models (LMMs). D…

Outlier Detection