paper-with-me

홈 › Papers

Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach

2026-02-26 · Jodi M. Casabianca, Maggie Beiting-Parrish arxiv

Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psychometric rater models into the AI pipeline to improve the reliability and validity of conclusions drawn from human judgments. The paper reviews common rater effects, severity and centrality, that distort observed ratings, and demonstrates how item response theory rater models, particularly the multi-faceted Rasch model, can separate true output quality from rater behavior. Using the OpenAI summarization dataset as an empirical example, we show how adjusting for rater severity produces corrected estimates of summary quality and provides diagnostic insight into rater performance. Incorporating psychometric modeling into human-in-the-loop evaluation offers more principled and transparent use of human data, enabling developers to make decisions based on adjusted scores rather than raw, error-prone ratings. This perspective highlights a path toward more robust, interpretable, and construct-aligned practices for AI development and evaluation.

📄 PDF Abstract BibTeX arXiv:2602.22585

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding and Reducing Crater Counting Errors in Citizen Science Data and the Need for Standardisation

2022-09-06 · P. D. Tar, N. A. Thacker

Citizen science has become a popular tool for preliminary data processing tasks, such as identifying and counting Lunar impact craters in modern high-resolution imagery. However, use of such data requires that citizen sc…

Investigating AI Rater Effects of Large Language Models: GPT, Claude, Gemini, and DeepSeek

2025-05-24 · Hong Jiao, Dan Song, Won-Chan Lee

Large language models (LLMs) have been widely explored for automated scoring in low-stakes assessment to facilitate learning and instruction. Empirical evidence related to which LLM produces the most reliable scores and …

Rich Insights from Cheap Signals: Efficient Evaluations via Tensor Factorization

2026-03-02 · Felipe Maia Polo, Aida Nematzadeh, Virginia Aglietti, Adam Fisch 외 arxiv

Moving beyond evaluations that collapse performance across heterogeneous prompts toward fine-grained evaluation at the prompt level, or within relatively homogeneous subsets, is necessary to diagnose generative models' s…

ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation

2024-05-08 · Ana Brassard, Benjamin Heinzerling, Keito Kudo, Keisuke Sakaguchi 외

Evaluating the quality of free-text explanations is a multifaceted, subjective, and labor-intensive task. Large language models (LLMs) present an appealing alternative due to their potential for consistency, scalability,…

Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing

2025-07-02 · Inyoung Cheong, Alicia Guo, Mina Lee, Zhehui Liao 외 arxiv

As AI integrates in various types of human writing, calls for transparency around AI assistance are growing. However, if transparency operates on uneven ground and certain identity groups bear a heavier cost for being ho…