paper-with-me

Papers

CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment

2025-12-11 · Yakun Zhu, Zhongzhen Huang, Qianhan Feng, Linjie Mu, Yannian Gu, Shaoting Zhang, Qi Dou, Xiaofan Zhang arxiv

Medical care follows complex clinical pathways that extend beyond isolated physician-patient encounters, emphasizing decision-making and transitions between different stages. Current benchmarks focusing on static exams or isolated dialogues inadequately evaluate large language models (LLMs) in dynamic clinical scenarios. We introduce CP-Env, a controllable agentic hospital environment designed to evaluate LLMs across end-to-end clinical pathways. CP-Env simulates a hospital ecosystem with patient and physician agents, constructing scenarios ranging from triage and specialist consultation to diagnostic testing and multidisciplinary team meetings for agent interaction. Following real hospital adaptive flow of healthcare, it enables branching, long-horizon task execution. We propose a three-tiered evaluation framework encompassing Clinical Efficacy, Process Competency, and Professional Ethics. Results reveal that most models struggle with pathway complexity, exhibiting hallucinations and losing critical diagnostic details. Interestingly, excessive reasoning steps can sometimes prove counterproductive, while top models tend to exhibit reduced tool dependency through internalized knowledge. CP-Env advances medical AI agents development through comprehensive end-to-end clinical evaluation. We provide the benchmark and evaluation tools for further research and development at https://github.com/SPIRAL-MED/CP_ENV.

📄 PDF Abstract BibTeX arXiv:2512.10206

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Identifying Critical Pathways in Coronary Heart Disease via Fuzzy Subgraph Connectivity

2025-09-19 · Shanookha Ali, Nitha Niralda P C arxiv

Coronary heart disease (CHD) arises from complex interactions among uncontrollable factors, controllable lifestyle factors, and clinical indicators, where relationships are often uncertain. Fuzzy subgraph connectivity (F…

MAP: Evaluation and Multi-Agent Enhancement of Large Language Models for Inpatient Pathways

2025-03-17 · Zhen Chen, Zhihao Peng, Xusheng Liang, Cheng Wang 외

Inpatient pathways demand complex clinical decision-making based on comprehensive patient information, posing critical challenges for clinicians. Despite advancements in large language models (LLMs) in medical applicatio…

Decision MakingMedical Question AnsweringQuestion Answering

When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering

2026-05-20 · Doeun Lee, Muge Zhang, Yi Yu, Ashish Manne 외 arxiv

Across medical specialties, clinical practice is anchored in evidence-based guidelines that codify best studied diagnostic and treatment pathways. These pathways routinely fall short for the long tail of real-world care …

Question Answering

Unveiling the Potential of Diffusion Large Language Model in Controllable Generation

2025-07-06 · Zhen Xiong, Yujun Cai, Zhecheng Li, Yiwei Wang arxiv

Controllable generation is a fundamental task in NLP with many applications, providing a basis for function calling to agentic communication. However, even state-of-the-art autoregressive Large Language Models (LLMs) tod…

Agentic Driving Coach: Robustness and Determinism of Agentic AI-Powered Human-in-the-Loop Cyber-Physical Systems

2026-04-13 · Deeksha Prahlad, Daniel Fan, Hokeun Kim arxiv

Foundation models, including large language models (LLMs), are increasingly used for human-in-the-loop (HITL) cyber-physical systems (CPS) because foundation model-based AI agents can potentially interact with both the p…