paper-with-me

홈 › Papers

Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications

2026-06-16 · Divyansh Srivastava, Shreya Ghosh, Anshul Verma, Rajkumar Buyya arxiv

Recent advances in Large Language Models (LLMs) and multi-agent systems have driven the rise of Agentic AI, showing promise for medical reasoning. However, open-ended conversational agents remain prone to two critical failure modes: premature diagnostic handoff and silent clinical hallucinations that may go undetected before reaching the patient. In this work, we propose a multi-agent framework that addresses both issues by replacing ``LLM-as-a-judge'' routing with deterministic orchestration constraints. The framework incorporates two safety mechanisms. First, a neuro-symbolic state-tracking gate enforces completeness of the OLDCARTS clinical protocol (Onset, Location, Duration, Character, Aggravating/Alleviating factors, Radiation, Timing, and Severity) by blocking diagnostic transitions until all required dimensions are collected. Second, an epistemic uncertainty quantification (UQ) gate computes semantic entropy (H) across K=5 independent diagnostic samples to identify and intercept divergent outputs before delivery. We evaluate the system using simulated patient agents powered by the llama-3.1-70b-instruct model on 150 test cases. The full architecture achieves 49.3% diagnostic precision, representing an absolute improvement of 11.3 percentage points over an unconstrained baseline. Additionally, we observe a statistically significant negative correlation (r = -0.181, p < 0.05) between OLDCARTS completeness (σ) and semantic entropy (H), suggesting that structured information gathering is associated with reduced diagnostic uncertainty.

📄 PDF Abstract BibTeX arXiv:2606.18068

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Agentic Electronic Design Automation: A Handoff Perspective

2026-06-18 · Jiawei Liu, Peiyi Han, Yuntao Lu, Su Zheng 외 arxiv

Electronic design automation (EDA) is inherently multi-stage and handoff-heavy. Design artifacts, flow scripts, and engineering decisions cross tool, session, and organizational boundaries before final implementation, si…

Quantifying and Mitigating Premature Closure in Frontier LLMs

2026-05-14 · Rebecca Handler, Suhana Bedi, Nigam Shah arxiv

Premature closure, or committing to a conclusion before sufficient information is available, is a recognized contributor to diagnostic error but remains underexamined in large language models (LLMs). We define LLM premat…

HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

2026-06-04 · Lizhi Yang, Junheng Li, Nehar Poddar, Yiling Hou 외 arxiv

For a humanoid robot to be deployed in the real world, the choice of command space (i.e., the interface between task planning and whole-body control) is crucial. Existing whole-body controllers typically demand dense kin…

Beyond Task Success: Measuring Workflow Fidelity in LLM-Based Agentic Payment Systems

2026-05-07 · Donghao Huang, Joon Kiat Chua, Zhaoxia Wang arxiv

LLM-based multi-agent systems are increasingly deployed for payment workflows, yet prevailing metrics, Task Success Rate (TSR) and Agent Handoff F1-Score (HF1), capture only final outcomes or unordered routing decisions.…

Information-seeking failures of large language models in agentic clinical reasoning

2026-07-11 · Krischan Braitsch, Laura K. Schmalbrock, Theresa Weltermann, Andrew F. Berdel 외 arxiv

Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty. We developed an agentic evaluation framework in hematol…