paper-with-me

홈 › Papers

AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making

2026-06-02 · Sangwon Baek, Kyu Yeon Hur, Kyunga Kim arxiv

Clinical AI evaluation increasingly delegates scoring to large language models (LLMs) acting as AI raters, yet their scoring behavior across evaluation conditions has not been quantitatively characterized. We address this gap through a factorial study of AI rater behavior in adult type 2 diabetes (T2D) pharmacotherapy at 12-month outpatient follow-up, a clinical task involving complex decision-making operationalized across seven evaluation questions. Four open-source LLMs served simultaneously as clinical decision support system (CDSS) models and AI raters. Each CDSS output was scored under two scoring protocols: a rubric-anchored Gold Rubric (GR) protocol incorporating a patient-specific rubric, and a rubric-free Non Gold Rubric (Non-GR) protocol. Linear mixed effects models crossed the scoring protocol factor with five design factors -- CDSS model, CDSS prompt configuration (document-referenced generation [DRG] vs.\ Baseline), rater model, prompt character, and prompt type -- and estimated main effects together with their protocol interactions. Across all questions, AI raters yielded consistently higher scores within a very narrow range (74--78 points on average) under Non-GR compared to those under GR (7.69 to 49.64 points lower mean scores; 1.68 to 3.67 times wider interquartile ranges). Within each question, GR amplified the AI rater's discrimination between DRG and Baseline CDSS outputs by factors of 1.76 to 5.10, while also revealing substantial behavioral variation across rater models that Non-GR suppressed. These findings support rubric anchoring as the scoring protocol that preserves discriminative power in clinical AI evaluation; rubric-free scoring cannot substitute when questions require patient-specific or jurisdiction-specific criteria that rater models cannot infer from parametric knowledge alone.

📄 PDF Abstract BibTeX arXiv:2606.03198

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Classical Versus Deep Mirror-Symmetry Scoring: A Benchmark of Thirteen Methods

2026-07-09 · Maximilian Woehrer arxiv

Quantifying how mirror-symmetric an image is about a given axis (symmetry scoring) underpins applications from visual aesthetics to medical imaging, yet proposed scoring methods have never been compared on a common, stat…

Symmetry Detection

Automated Essay Scoring in the Presence of Biased Ratings

2018-06-01 · NAACL 2018 6 · Evelin Amorim, Marcia Can{\c{c}}ado, Adriano Veloso

Studies in Social Sciences have revealed that when people evaluate someone else, their evaluations often reflect their biases. As a result, rater bias may introduce highly subjective factors that make their evaluations i…

Automated Essay Scoring

Investigating AI Rater Effects of Large Language Models: GPT, Claude, Gemini, and DeepSeek

2025-05-24 · Hong Jiao, Dan Song, Won-Chan Lee

Large language models (LLMs) have been widely explored for automated scoring in low-stakes assessment to facilitate learning and instruction. Empirical evidence related to which LLM produces the most reliable scores and …

Using Large Language Models to Assess Teachers' Pedagogical Content Knowledge

2025-05-25 · Yaxuan Yang, Shiyu Wang, Xiaoming Zhai

Assessing teachers' pedagogical content knowledge (PCK) through performance-based tasks is both time and effort-consuming. While large language models (LLMs) offer new opportunities for efficient automatic scoring, littl…

Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement

2026-05-07 · Jessica Huynh, Alfredo Gomez, Athiya Deviyani, Renee Shelby 외 arxiv

Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifications in a rubric presented to both huma…