paper-with-me

홈 › Papers

MedConceal: A Benchmark for Clinical Hidden-Concern Reasoning Under Partial Observability

2026-04-09 · Yikun Han, Joey Chan, Jingyuan Chen, Mengting Ai, Simo Du, Yue Guo arxiv

Patient-clinician communication is an asymmetric-information problem: patients often do not disclose fears, misconceptions, or practical barriers unless clinicians elicit them skillfully. Effective medical dialogue therefore requires reasoning under partial observability: clinicians must elicit latent concerns, confirm them through interaction, and respond in ways that guide patients toward appropriate care. However, existing medical dialogue benchmarks largely sidestep this challenge by exposing hidden patient state, collapsing elicitation into extraction, or evaluating responses without modeling what remains hidden. We present MedConceal, a benchmark with an interactive patient simulator for evaluating hidden-concern reasoning in medical dialogue, comprising 300 curated cases and 600 clinician-LLM interactions. Built from clinician-answered online health discussions, each case pairing clinician-visible context with simulator-internal hidden concerns derived from prior literature and structured using an expert-developed taxonomy. The simulator withholds these concerns from the dialogue agent, tracks whether they have been revealed and addressed via theory-grounded turn-level communication signals, and is clinician-reviewed for clinical plausibility. This enables process-aware evaluation of both task success and the interaction process that leads to it. We study two abilities: confirmation, surfacing hidden concerns through multi-turn dialogue, and intervention, addressing the primary concern and guiding the patient toward a target plan. Results show that no single system dominates: frontier models lead on different confirmation metrics, while human clinicians (N=159) remain strongest on intervention success. Together, these results identify hidden-concern reasoning under partial observability as a key unresolved challenge for medical dialogue systems.

📄 PDF Abstract BibTeX arXiv:2604.08788

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

2026-05-11 · Minh Khoi Nguyen, Dai Lam Le, Amir Reza Jafari, Tuan Dung Nguyen 외 arxiv

Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet incorrect statements, raising significant safety concerns. Existing medi…

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

2026-06-30 · Ajmal M., Abin Roy, Afthab Salam Kanniyan, Jawadh Abdul Kabeer 외 arxiv

Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured ex…

Limitations of Large Language Models in Clinical Problem-Solving Arising from Inflexible Reasoning

2025-02-05 · Jonathan Kim, Anna Podlasek, Kie Shidara, Feng Liu 외

Large Language Models (LLMs) have attained human-level accuracy on medical question-answer (QA) benchmarks. However, their limitations in navigating open-ended clinical scenarios have recently been shown, raising concern…

ARC

Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation

2025-07-29 · Ziyao Wang, Guoheng Sun, Yexiao He, Zheyu Shen 외 arxiv

Commercial LLM services often conceal internal reasoning traces while still charging users for every generated token, including those from hidden intermediate steps, raising concerns of token inflation and potential over…

The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1

2025-02-18 · Kaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Shreedhar Jangam 외

The rapid development of large reasoning models, such as OpenAI-o3 and DeepSeek-R1, has led to significant improvements in complex reasoning over non-reasoning large language models~(LLMs). However, their enhanced capabi…