paper-with-me

홈 › Papers

Two Axes of LLM Abstention: Answer Correctness and Question Answerability

2026-07-09 · Benedikt J. Wagner arxiv

A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable ones or ones resting on a false premise. The usual recipe thresholds a single confidence score, which cannot tell these apart. Across five instruction-tuned models from three families (2B to 14B), we find they are separate axes. Ordinary answer-confidence tracks whether an answer is right but is nearly blind to whether the question is answerable; a linear probe on hidden states does the reverse. The blind spot does not shrink with scale. It is worst on naturally occurring false-premise questions (CREPE). There, answer-confidence, P(IK), P(True), and even asking the model outright whether a premise is false all stay near chance, while a hidden-state probe reaches 0.69 to 0.77 AUROC: the model represents a problem it will not report. This turns out to be fixable. Instructing a model to check premises backfires, because it then disputes sound and false premises alike (57% false challenges), unable to tell them apart; routing the same instruction with the probe roughly triples challenge precision. We turn the two axes into a calibrated policy that answers only when an answerability score and a correctness score each clear a separately certifies behave differently: the unanswerable-answer rate is controllable at every scale, while the wrong-answer rate is capped by model accuracy, so the guarantee tightens as threshold policy certifies both budgets at 0.75 coverage of correct answers, against 0.31 for a single threshold; at 14B it is the only policy that certifies at all.

📄 PDF Abstract BibTeX arXiv:2607.08456

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems

2026-04-16 · Nishanth Madhusudhan, Vikas Yadav, Alexandre Lacoste arxiv

Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing evaluation paradigms for vision-language models (VLMs) and multi-agen…

Multimodal Reasoning

Detecting (Un)answerability in Large Language Models with Linear Directions

2025-09-26 · Maor Juliet Lavi, Tova Milo, Mor Geva arxiv

Large language models (LLMs) often respond confidently to questions even when they lack the necessary information, leading to hallucinated answers. In this work, we study the problem of (un)answerability detection, focus…

Question Answering

Harnessing RLHF for Robust Unanswerability Recognition and Trustworthy Response Generation in LLMs

2025-07-22 · Shuyuan Lin, Lei Duan, Philip Hughes, Yuxuan Sheng arxiv

Conversational Information Retrieval (CIR) systems, while offering intuitive access to information, face a significant challenge: reliably handling unanswerable questions to prevent the generation of misleading or halluc…

Reinforcement LearningInformation RetrievalMulti-Task LearningResponse Generation

Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information

2026-05-27 · Renjie Gu, Jiaxu Li, Yihao Wang, Yun Yue 외 arxiv

We highlight a failure mode of large reasoning models on questions with insufficient information: models may recognize that a problem is under-specified, yet still continue reasoning and produce unsupported final answers…

Reinforcement Learning

Reported Confidence in LLMs Tracks Commitment More Than Correctness

2026-06-28 · Dharshan Kumaran arxiv

Confidence is an estimate of the probability that a chosen answer is correct. Verbal confidence reports are widely used as uncertainty measures in large language models, but whether they are best understood as estimates …

Decision Making