paper-with-me

홈 › Papers

Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation

2026-06-14 · Gili Lior, Tzviel Frostig, Gabriel Stanovsky, Matan Eyal arxiv

Multilingual benchmarks are central to evaluating large language models (LLMs) across languages, but they suffer from three issues: exhaustive evaluation scales linearly with the number of languages, automatic translation introduces errors that are easily missed at scale, and some items conflate general and culture-specific knowledge. We address all three with a unified statistical framework, Multilingual-IRT, which extends Item Response Theory with per-language difficulty deviations, split discriminability separating content from language effects, and per-language ability residuals. Fitting Multilingual-IRT on 25 LLMs across 29 languages of MMLU-Pro-X, we show that its fitted parameters support three practical applications: predicting unobserved (item, LLM, language) instances with 11-16% lower binary cross-entropy than the strongest accuracy-based baseline, surfacing candidate translation errors distributed across all 28 non-English languages, whereas accuracy-based baselines concentrate detections in a few languages, and recovering culture-specific items that accuracy-based baselines miss.

📄 PDF Abstract BibTeX arXiv:2606.15643

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Estimating Heterogeneous Treatment Effects with Item-Level Outcome Data: Insights from Item Response Theory

2024-04-30 · Joshua B. Gilbert, Zachary Himmelsbach, James Soland, Mridul Joshi 외

Analyses of heterogeneous treatment effects (HTE) are common in applied causal inference research. However, when outcomes are latent variables assessed via psychometric instruments such as educational tests, standard met…

Causal Inference

Capturing Humans' Mental Models of AI: An Item Response Theory Approach

2023-05-15 · Markelle Kelly, Aakriti Kumar, Padhraic Smyth, Mark Steyvers

Improving our understanding of how humans perceive AI teammates is an important foundation for our general understanding of human-AI teams. Extending relevant work from cognitive science, we propose a framework based on …

AI AgentQuestion Answering

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

2026-06-17 · Han Chen, Ming Li, Chenguang Wang, Yijun Liang 외 arxiv

Item discrimination is a fundamental psychometric property of educational assessment, which measures whether an item meaningfully distinguishes students with higher proficiency from students with lower proficiency. While…

Reading Comprehension

TopicResponse: A Marriage of Topic Modelling and Rasch Modelling for Automatic Measurement in MOOCs

2016-07-29 · Jiazhen He, Benjamin I. P. Rubinstein, James Bailey, Rui Zhang 외

This paper explores the suitability of using automatically discovered topics from MOOC discussion forums for modelling students' academic abilities. The Rasch model from psychometrics is a popular generative probabilisti…

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

2026-07-06 · Andrei Florian, Cynthia Jayne Amol, Hope Kerubo Ombaba, Xiaoyu Cui 외 arxiv

Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale. A promising direction is to leverage linguistic relatedness to enhance…

Cross-Lingual TransferSpeech Recognition