paper-with-me

홈 › Papers

Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis

2025-12-16 · Hongli Li, Che Han Chen, Kevin Fan, Chiho Young-Johnson, Soyoung Lim, Yali Feng arxiv

Despite the growing promise of large language models (LLMs) in automated essay scoring (AES), empirical findings regarding their reliability compared to human raters remain mixed. Following the PRISMA 2020 guidelines, we synthesized 65 published and unpublished studies from January 2022 to August 2025 that examined agreement between LLM-generated scores and human ratings. Agreement levels varied substantially both across and within studies, with reported values spanning a wide range. Overall, the findings suggest that LLM-human agreement is highly context-dependent. Implications, challenges, and directions for future research are discussed.

📄 PDF Abstract BibTeX arXiv:2512.14561

Code (0)

등록된 구현이 없습니다.

Tasks

Automated Essay Scoring

Similar Papers 제목 키워드 기반

Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals

2026-05-12 · Yo Ehara arxiv

Automatic generation of educational materials using large language models (LLMs) is becoming increasingly common, but assigning difficulty levels to such materials still requires substantial human effort. LLM-as-a-Judge …

Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters

2025-10-29 · Xingjian Zhang, Tianhong Gao, Suliang Jin, Tianhao Wang 외 arxiv

Large language models (LLMs) are increasingly used as raters for evaluation tasks. However, their reliability is often limited for subjective tasks, when human judgments involve subtle reasoning beyond annotation labels.…

Calibrating Generative AI to Produce Realistic Essays for Data Augmentation

2026-02-06 · Edward W. Wolfe, Justin O. Barber arxiv

Data augmentation can mitigate limited training data in machine-learning automated scoring engines for constructed response items. This study seeks to determine how well three approaches to large language model prompting…

Data Augmentation

ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation

2024-05-08 · Ana Brassard, Benjamin Heinzerling, Keito Kudo, Keisuke Sakaguchi 외

Evaluating the quality of free-text explanations is a multifaceted, subjective, and labor-intensive task. Large language models (LLMs) present an appealing alternative due to their potential for consistency, scalability,…

Rater Cohesion and Quality from a Vicarious Perspective

2024-08-15 · Deepak Pandita, Tharindu Cyril Weerasooriya, Sujan Dutta, Sarah K. Luger 외

Human feedback is essential for building human-centered AI systems across domains where disagreement is prevalent, such as AI safety, content moderation, or sentiment analysis. Many disagreements, particularly in politic…

Sentiment Analysis