paper-with-me

홈 › Papers

Decomposing Physician Disagreement in HealthBench

2026-02-26 · Satya Borgohain, Roy Mariathas arxiv

We decompose physician disagreement in the HealthBench medical AI evaluation dataset to understand where variance resides and what observable features can explain it. Rubric identity accounts for 15.8% of met/not-met label variance but only 3.6-6.9% of disagreement variance; physician identity accounts for just 2.4%. The dominant 81.8% case-level residual is not reduced by HealthBench's metadata labels (z = -0.22, p = 0.83), normative rubric language (pseudo R^2 = 1.2%), medical specialty (0/300 Tukey pairs significant), surface-feature triage (AUC = 0.58), or embeddings (AUC = 0.485). Disagreement follows an inverted-U with completion quality (AUC = 0.689), confirming physicians agree on clearly good or bad outputs but split on borderline cases. Physician-validated uncertainty categories reveal that reducible uncertainty (missing context, ambiguous phrasing) more than doubles disagreement odds (OR = 2.55, p < 10^(-24)), while irreducible uncertainty (genuine medical ambiguity) has no effect (OR = 1.01, p = 0.90), though even the former explains only ~3% of total variance. The agreement ceiling in medical AI evaluation is thus largely structural, but the reducible/irreducible dissociation suggests that closing information gaps in evaluation scenarios could lower disagreement where inherent clinical ambiguity does not, pointing toward actionable evaluation design improvements.

📄 PDF Abstract BibTeX arXiv:2602.22758

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HealthBench: Evaluating Large Language Models Towards Improved Human Health

2025-05-13 · Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman 외

We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare. HealthBench consists of 5,000 multi-turn conversations between a model and an individual user …

Instruction FollowingMultiple-choice

HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats

2026-04-30 · Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim, Rahul K. Arora 외 arxiv

Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Professional, an open benchmark for evaluat…

Rethinking Evidence Hierarchies in Medical Language Benchmarks: A Critical Evaluation of HealthBench

2025-07-31 · Fred Mutisya, Shikoh Gitau, Nasubo Ongoma, Keith Mbae 외 arxiv

HealthBench, a benchmark designed to measure the capabilities of AI systems for health better (Arora et al., 2025), has advanced medical language model evaluation through physician-crafted dialogues and transparent rubri…

Reinforcement Learning

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

2026-06-27 · Jean Feng, Vishal Patel, Patrick Heagerty, Yifan Mai 외 arxiv

Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions, not those actually asked in practice. We report a blinded evaluatio…

Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making

2026-02-06 · Baichuan-M3 Team, :, Chengfeng Dou, Fan Yang 외 arxiv

We introduce Baichuan-M3, a medical-enhanced large language model engineered to shift the paradigm from passive question-answering to active, clinical-grade decision support. Addressing the limitations of existing system…