paper-with-me

홈 › Papers

Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis

2026-01-27 · Zhuohan Long, Zhijie Bao, Zhongyu Wei arxiv

Interactive medical consultation requires an agent to proactively elicit missing clinical evidence under uncertainty. Yet existing evaluations largely remain static or outcome-centric, neglecting the evidence-gathering process. In this work, we propose an interactive evaluation framework that explicitly models the consultation process using a simulated patient and a \rev{simulated reporter} grounded in atomic evidences. Based on this representation, we introduce Information Coverage Rate (ICR) to quantify how completely an agent uncovers necessary evidence during interaction. To support systematic study, we build EviMed, an evidence-based benchmark spanning diverse conditions from common complaints to rare diseases, and evaluate 10 models with varying reasoning abilities. We find that strong diagnostic reasoning does not guarantee effective information collection, and this insufficiency acts as a primary bottleneck limiting performance in interactive settings. To address this, we propose REFINE, a strategy that leverages diagnostic verification to guide the agent in proactively resolving uncertainties. Extensive experiments demonstrate that REFINE consistently outperforms baselines across diverse datasets and facilitates effective model collaboration, enabling smaller agents to achieve superior performance under strong reasoning supervision. Our code can be found at https://github.com/NanshineLoong/EID-Benchmark .

📄 PDF Abstract BibTeX arXiv:2601.19773

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs

2025-10-17 · Hongcheng Liu, Pingjie Wang, Yuhao Wang, Siqu Ou 외 arxiv

Multimodal large language models (MLLMs) have shown strong capabilities across a broad range of benchmarks. However, most existing evaluations focus on passive inference, where models perform step-by-step reasoning under…

Uncertainty in Action: Confidence Elicitation in Embodied Agents

2025-03-13 · Tianjiao Yu, Vedant Shah, Muntasir Wahed, Kiet A. Nguyen 외

Expressing confidence is challenging for embodied agents navigating dynamic multimodal environments, where uncertainty arises from both perception and decision-making processes. We present the first work investigating em…

Decision MakingMinecraft

NeuReasoner: Theory-grounded Mapping of Reasoning Elicitation Boundaries

2026-06-29 · Aydin Javadov, Shyngys Aitkazinov, Tobias Hoesli, Florian von Wangenheim 외 arxiv

A growing body of work suggests that the reasoning capabilities of large language models are largely latent in their base form, with post-training primarily amplifying rather than introducing them. However, this evidence…

Arithmetic ReasoningCode GenerationDecision Making

Scalable Delphi: Large Language Models for Structured Risk Estimation

2026-02-09 · Tobias Lorenz, Mario Fritz arxiv

Quantitative risk assessment in high-stakes domains relies on structured expert elicitation to estimate unobservable properties. The gold standard - the Delphi method - produces calibrated, auditable judgments but requir…

Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding

2026-08-06 · Soojin Yoon, Dongha Lee arxiv

User requests serve as research specifications for deep research agents, shaping what evidence to seek and how to synthesize it. In personalized deep research, these specifications must additionally reflect user goals, c…