paper-with-me

홈 › Papers

Quantifying and Mitigating Premature Closure in Frontier LLMs

2026-05-14 · Rebecca Handler, Suhana Bedi, Nigam Shah arxiv

Premature closure, or committing to a conclusion before sufficient information is available, is a recognized contributor to diagnostic error but remains underexamined in large language models (LLMs). We define LLM premature closure as inappropriate commitment under uncertainty: providing an answer, recommendation, or clinical guidance when the safer response would be clarification, abstention, escalation, or refusal. We evaluated five frontier LLMs across structured and open-ended medical tasks. In MedQA (n = 500) and AfriMed-QA (n = 490) questions where the correct choice had been removed, models still selected an answer at high rates, with baseline false-action rates of 55-81% and 53-82%, respectively. In open-ended evaluation, models gave inappropriate answers on an average of 30% of 861 HealthBench questions and 78% of 191 physician-authored adversarial queries. Safety-oriented prompting reduced premature closure across models, but residual failure persisted, highlighting the need to evaluate whether medical LLMs know when not to answer.

📄 PDF Abstract BibTeX arXiv:2605.15000

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Operationalizing Data Minimization for Privacy-Preserving LLM Prompting

2025-10-04 · Jijie Zhou, Niloofar Mireshghallah, Tianshi Li arxiv

The rapid deployment of large language models (LLMs) in consumer applications has led to frequent exchanges of personal information. To obtain useful responses, users often share more than necessary, increasing privacy r…

TopoBench: Benchmarking LLMs on Hard Topological Reasoning

2026-03-12 · Mayug Maniparambil, Nils Hoehing, Janak Kapuriya, Arjun Karuvally 외 arxiv

Solving topological grid puzzles requires reasoning over global spatial invariants such as connectivity, loop closure, and region symmetry and remains challenging for even the most powerful large language models (LLMs). …

What Did They Mean? How LLMs Resolve Ambiguous Social Situations across Perspectives and Roles

2026-04-27 · Qiming Yuan, Linyi Han, Nam Ling, Cihan Ruan arxiv

People increasingly turn to large language models (LLMs) to interpret ambiguous social situations: a delayed text reply, an unusually cold supervisor, a teacher's mixed signals, or a boundary-crossing friend. Yet in many…

Beyond Performance: Quantifying and Mitigating Label Bias in LLMs

2024-05-04 · Yuval Reif, Roy Schwartz

Large language models (LLMs) have shown remarkable adaptability to diverse tasks, by leveraging context prompts containing instructions, or minimal input-output examples. However, recent work revealed they also exhibit l…

Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure

2026-03-17 · Caglar Yildirim arxiv

Large language models (LLMs) are increasingly deployed as tool-using agents, shifting safety concerns from harmful text generation to harmful task completion. Deployed systems often condition on user profiles or persiste…

Text Generation