paper-with-me

홈 › Papers

Active Evidence-Seeking and Diagnostic Reasoning in Large Language Models for Clinical Decision Support

2026-05-21 · Chen Zhan, Xihe Qiu, Xiaoyu Tan, Xibing Zhuang, Gengchen Ma, Yue Zhang, Shuo Li, Peifeng Liu, Xiaoxiao Ge, Liang Liu, Lu Gan arxiv

Large language models perform well on static medical examinations, yet clinical diagnosis often requires iterative evidence gathering under uncertainty. Building on prior interactive evaluation efforts, we introduce an OSCE-inspired standardized patient simulator and a controlled, reproducible benchmark for active diagnostic inquiry. Across 468 cases and 15 models in our protocol, we observe that multi-turn evidence seeking reduces diagnostic accuracy by 12.75% and lowers supporting-evidence quality by 24.36% relative to full-context evaluation; error analyses associate these drops with premature diagnostic closure and inefficient questioning. Together, these results suggest that static full-context benchmarks may overestimate performance in interactive evidence-seeking settings, motivating complementary interactive assessment for safer clinical decision support.

📄 PDF Abstract BibTeX arXiv:2605.22047

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis

2025-12-29 · Shengyi Hua, Jianfeng Wu, Tianle Shen, Kangzhe Hu 외 arxiv

Recent pathological foundation models have substantially advanced visual representation learning and multimodal interaction. However, most models still rely on a static inference paradigm in which whole-slide images are …

Representation LearningReinforcement Learning

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

2026-07-21 · Dankai Liao, Tianyi Zhang, Yufeng Wu, Xinyue Zhang 외 arxiv

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scale evidence. However, most existing pathology benchmarks evaluate mod…

Image Retrieval

Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models

2026-07-03 · Shengyi Hua, Kangzhe Hu, Conghui He, Xiaofan Zhang 외 arxiv

Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that assumes complete information. In contrast, real-world clinical intel…

Reinforcement LearningMedical Diagnosis

MedClarify: An information-seeking AI agent for medical diagnosis with case-specific follow-up questions

2026-02-19 · Hui Min Wong, Philip Heesen, Pascal Janetzky, Martin Bendszus 외 arxiv

Large language models (LLMs) are increasingly used for diagnostic tasks in medicine. In clinical practice, the correct diagnosis can rarely be immediately inferred from the initial patient presentation alone. Rather, rea…

Medical Diagnosis

UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs

2026-05-12 · Shuo Ni, Tong Wang, Jing Zhang, He Chen 외 arxiv

Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scale mismatch between large-scale scene context and micro-scale targets…