paper-with-me

홈 › Papers

CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine

2026-05-01 · Kevin H. Guo, Chao Yan, Avinash Baidya, Katherine Brown, Xiang Gao, Juming Xiong, Zhijun Yin, Bradley A. Malin arxiv

Medical large language model (LLM) evaluations rely on simplified, exam-style benchmarks that rarely reflect the ambiguity of real-world medical inquiries. We introduce the CLinical Evaluation of Ambiguity and Reliability (CLEAR) framework, which assesses how decision-space presentation, ambiguity, and uncertainty affect LLMs' reasoning on medical benchmarks. CLEAR systematically perturbs (1) the number of plausible answer options, (2) the presence of a ground truth or abstention option, and (3) the semantic framing of answer options. Applying CLEAR on three benchmarks evaluated across 17 LLMs reveals three notable limitations of existing evaluation methods. First, increasing the number of plausible answers degrades a model's ability to identify the correct answer and abstain against incorrect ones. Second, this lack of caution intensifies as the framing of abstention shifts from assertive rejection like "None of the Above" to uncertainty admission like "I don't know" (IDK). Notably, just including IDK in the answer space increases incorrect answer selections. Lastly, we formalize the performance gap between identifying the correct answer and abstaining from incorrect ones as the humility deficit, which worsens with model scale. Our findings reveal limitations in standard medical benchmarks and underscore that scaling alone does not resolve LLM reliability issues.

📄 PDF Abstract BibTeX arXiv:2605.01011

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Needle-Match: Reliable Patch Matching Under High Uncertainty

2016-06-01 · CVPR 2016 6 · Or Lotan, Michal Irani

Reliable patch-matching forms the basis for many algorithms (super-resolution, denoising, inpainting, etc.) However, when the image quality deteriorates (by noise, blur or geometric distortions), the reliability of patc…

DenoisingPatch MatchingSuper-ResolutionVocal Bursts Intensity Prediction

AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity

2025-09-17 · Yifan Liu, Wenkuan Zhao, Shanshan Zhong, Jinghui Qin 외 arxiv

Recent advancements in multimodal large language models (MLLMs) have garnered significant attention, offering a promising pathway toward artificial general intelligence (AGI). Among the essential capabilities required fo…

The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity

2025-11-06 · Tim Tomov, Dominik Fuchsgruber, Tom Wollschläger, Stephan Günnemann arxiv

Accurate uncertainty quantification (UQ) in Large Language Models (LLMs) is critical for trustworthy deployment. While real-world language is inherently ambiguous, reflecting aleatoric uncertainty, existing UQ methods ar…

NightHaze: Nighttime Image Dehazing via Self-Prior Learning

2024-03-12 · Beibei Lin, Yeying Jin, Wending Yan, Wei Ye 외

Masked autoencoder (MAE) shows that severe augmentation during training produces robust representations for high-level tasks. This paper brings the MAE-like framework to nighttime image enhancement, demonstrating that se…

Image DehazingImage Enhancement

SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation

2026-03-31 · Ryosuke Matsuda, Keito Kudo, Haruto Yoshida, Nobuyuki Shimizu 외 arxiv

This paper proposes the synthetic long-video meta-evaluation (SLVMEval), a benchmark for meta-evaluating text-to-video (T2V) evaluation systems. The proposed SLVMEval benchmark focuses on assessing these systems on video…

Video Generation