paper-with-me

홈 › Papers

MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences

2026-03-13 · Eric Wu, Kevin Wu, Jason Hom, Paul H. Yi, Angela Zhang, Alejandro Lozano, Jeff Nirschl, Jeff Tangney, Kevin Byram, Braydon Dymm, Narender Annapureddy, Eric Topol, David Ouyang, James Zou arxiv

Large language models (LLMs) are increasingly central to clinician workflows, spanning clinical decision support, medical education, and patient communication. However, current evaluation methods for medical LLMs rely heavily on static, templated benchmarks that fail to capture the complexity and dynamics of real-world clinical practice, creating a dissonance between benchmark performance and clinical utility. To address these limitations, we present MedArena, an interactive evaluation platform that enables clinicians to directly test and compare leading LLMs using their own medical queries. Given a clinician-provided query, MedArena presents responses from two randomly selected models and asks the user to select the preferred response. Out of 1571 preferences collected across 12 LLMs up to November 1, 2025, Gemini 2.0 Flash Thinking, Gemini 2.5 Pro, and GPT-4o were the top three models by Bradley-Terry rating. Only one-third of clinician-submitted questions resembled factual recall tasks (e.g., MedQA), whereas the majority addressed topics such as treatment selection, clinical documentation, or patient communication, with ~20% involving multi-turn conversations. Additionally, clinicians cited depth and detail and clarity of presentation more often than raw factual accuracy when explaining their preferences, highlighting the importance of readability and clinical nuance. We also confirm that the model rankings remain stable even after controlling for style-related factors like response length and formatting. By grounding evaluation in real-world clinical questions and preferences, MedArena offers a scalable platform for measuring and improving the utility and efficacy of medical LLMs.

📄 PDF Abstract BibTeX arXiv:2603.15677

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

2026-05-07 · Jinge Wu, Hongjian Zhou, Mingde Zeng, Jiayuan Zhu 외 arxiv

Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness and tool registry differ, and integrating …

Diagnostic Reasoning Prompts Reveal the Potential for Large Language Model Interpretability in Medicine

2023-08-13 · Thomas Savage, Ashwin Nayak, Robert Gallo, Ekanath Rangan 외

One of the major barriers to using large language models (LLMs) in medicine is the perception they use uninterpretable methods to make clinical decisions that are inherently different from the cognitive processes of clin…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model

Differentiating hype from practical applications of large language models in medicine -- a primer for healthcare professionals

2025-07-25 · Elisha D. O. Roberson arxiv

The medical ecosystem consists of the training of new clinicians and researchers, the practice of clinical medicine, and areas of adjacent research. There are many aspects of these domains that could benefit from the app…

Towards Accurate Differential Diagnosis with Large Language Models

2023-11-30 · Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu 외

An accurate differential diagnosis (DDx) is a cornerstone of medical care, often reached through an iterative process of interpretation that combines clinical history, physical examination, investigations and procedures.…

Diagnostic

Speech Detection For Child-Clinician Conversations In Danish For Low-Resource In-The-Wild Conditions: A Case Study

2022-04-25 · Sneha Das, Nicole Nadine Lønfeldt, Anne Katrine Pagsberg, Line. H. Clemmensen

Use of speech models for automatic speech processing tasks can improve efficiency in the screening, analysis, diagnosis and treatment in medicine and psychiatry. However, the performance of pre-processing speech tasks li…

Classification