paper-with-me

홈 › Papers

Moving Beyond Medical Exam Questions: A Clinician-Annotated Dataset of Real-World Tasks and Ambiguity in Mental Healthcare

2025-02-22 · Max Lamparth, Declan Grabb, Amy Franks, Scott Gershan, Kaitlyn N. Kunstman, Aaron Lulla, Monika Drummond Roots, Manu Sharma, Aryan Shrivastava, Nina Vasan, Colleen Waickman

Current medical language model (LM) benchmarks often over-simplify the complexities of day-to-day clinical practice tasks and instead rely on evaluating LMs on multiple-choice board exam questions. Thus, we present an expert-created and annotated dataset spanning five critical domains of decision-making in mental healthcare: treatment, diagnosis, documentation, monitoring, and triage. This dataset - created without any LM assistance - is designed to capture the nuanced clinical reasoning and daily ambiguities mental health practitioners encounter, reflecting the inherent complexities of care delivery that are missing from existing datasets. Almost all 203 base questions with five answer options each have had the decision-irrelevant demographic patient information removed and replaced with variables (e.g., AGE), and are available for male, female, or non-binary-coded patients. For question categories dealing with ambiguity and multiple valid answer options, we create a preference dataset with uncertainties from the expert annotations. We outline a series of intended use cases and demonstrate the usability of our dataset by evaluating eleven off-the-shelf and four mental health fine-tuned LMs on category-specific task accuracy, on the impact of patient demographic information on decision-making, and how consistently free-form responses deviate from human annotated samples.

📄 PDF Abstract BibTeX arXiv:2502.16051

Code (1)

maxlampe/mentat 공식 구현

Tasks

Decision MakingMultiple-choice

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

EvoClinician: A Self-Evolving Agent for Multi-Turn Medical Diagnosis via Test-Time Evolutionary Learning

2026-01-30 · Yufei He, Juncheng Liu, Zhiyuan Hu, Yulin Chen 외 arxiv

Prevailing medical AI operates on an unrealistic ''one-shot'' model, diagnosing from a complete patient file. However, real-world diagnosis is an iterative inquiry where Clinicians sequentially ask questions and order te…

Continual LearningMedical Diagnosis

Language models are susceptible to incorrect patient self-diagnosis in medical applications

2023-09-17 · Rojin Ziaei, Samuel Schmidgall

Large language models (LLMs) are becoming increasingly relevant as a potential tool for healthcare, aiding communication between clinicians, researchers, and patients. However, traditional evaluations of LLMs on medical …

DiagnosticMultiple-choice

PubMed and Beyond: Biomedical Literature Search in the Age of Artificial Intelligence

2023-07-18 · Qiao Jin, Robert Leaman, Zhiyong Lu

Biomedical research yields a wealth of information, much of which is only accessible through the literature. Consequently, literature search is an essential tool for building on prior knowledge in clinical and biomedical…

ArticlesSurvey

MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences

2026-03-13 · Eric Wu, Kevin Wu, Jason Hom, Paul H. Yi 외 arxiv

Large language models (LLMs) are increasingly central to clinician workflows, spanning clinical decision support, medical education, and patient communication. However, current evaluation methods for medical LLMs rely he…

ClinAlign: Scaling Healthcare Alignment from Clinician Preference

2026-02-10 · Shiwei Lyu, Xidong Wang, Lei Liu, Hao Zhu 외 arxiv

Although large language models (LLMs) demonstrate expert-level medical knowledge, aligning their open-ended outputs with fine-grained clinician preferences remains challenging. Existing methods often rely on coarse objec…