paper-with-me

홈 › Papers

Item Response Theory for Efficient Human Evaluation of Chatbots

2020-11-01 · EMNLP (Eval4NLP) 2020 11 · João Sedoc, Lyle Ungar

Conversational agent quality is currently assessed using human evaluation, and often requires an exorbitant number of comparisons to achieve statistical significance. In this paper, we introduce Item Response Theory (IRT) for chatbot evaluation, using a paired comparison in which annotators judge which system responds better to the next turn of a conversation. IRT is widely used in educational testing for simultaneously assessing the ability of test takers and the quality of test questions. It is similarly well suited for chatbot evaluation since it allows the assessment of both models and the prompts used to evaluate them. We use IRT to efficiently assess chatbots, and show that different examples from the evaluation set are better suited for comparing high-quality (nearer to human performance) than low-quality systems. Finally, we use IRT to reduce the number of evaluation examples assessed by human annotators while retaining discriminative power.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Chatbot

Similar Papers 제목 키워드 기반

Assessment Design in the AI Era: A Method for Identifying Items Functioning Differentially for Humans and Chatbots

2026-03-24 · Licol Zeinfeld, Alona Strugatski, Ziva Bar-Dov, Ron Blonder 외 arxiv

The rapid adoption of large language models (LLMs) in education raises profound challenges for assessment design. To adapt assessments to the presence of LLM-based tools, it is crucial to characterize the strengths and w…

Applying IRT to Distinguish Between Human and Generative AI Responses to Multiple-Choice Assessments

2024-11-28 · Alona Strugatski, Giora Alexandron

Generative AI is transforming the educational landscape, raising significant concerns about cheating. Despite the widespread use of multiple-choice questions in assessments, the detection of AI cheating in MCQ-based test…

Multiple-choice

Building an Evaluation Scale using Item Response Theory

2016-05-28 · EMNLP 2016 11 · John P. Lalor, Hao Wu, Hong Yu

Evaluation of NLP methods requires testing against a previously vetted gold-standard test set and reporting standard metrics (accuracy/precision/recall/F1). The current assumption is that all items in a given test set ar…

Natural Language Inference

Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory

2025-05-21 · Hongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han 외

The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately refl…

BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model

Correcting Human Labels for Rater Effects in AI Evaluation: An Item Response Theory Approach

2026-02-26 · Jodi M. Casabianca, Maggie Beiting-Parrish arxiv

Human evaluations play a central role in training and assessing AI models, yet these data are rarely treated as measurements subject to systematic error. This paper integrates psychometric rater models into the AI pipeli…