paper-with-me

홈 › Papers

Hallucination as output-boundary misclassification: a composite abstention architecture for language models

2026-03-12 · Angelina Hintsanen arxiv

Large language models often produce unsupported claims. We frame this as a misclassification error at the output boundary, where internally generated completions are emitted as if they were grounded in evidence. This motivates a composite intervention that combines instruction-based refusal with a structural abstention gate. The gate computes a support deficit score, St, from three black-box signals: self-consistency (At), paraphrase stability (Pt), and citation coverage (Ct), and blocks output when St exceeds a threshold. In a controlled evaluation across 50 items, five epistemic regimes, and three models, neither mechanism alone was sufficient. Instruction-only prompting reduced hallucination sharply, but still showed over-cautious abstention on answerable items and residual hallucination for GPT-3.5-turbo. The structural gate preserved answerable accuracy across models but missed confident confabulation on conflicting-evidence items. The composite architecture achieved high overall accuracy with low hallucination, while also inheriting some over-abstention from the instruction component. A supplementary 100-item no-context stress test derived from TruthfulQA showed that structural gating provides a capability-independent abstention floor. Overall, instruction-based refusal and structural gating show complementary failure modes, which suggests that effective hallucination control benefits from combining both mechanisms.

📄 PDF Abstract BibTeX arXiv:2604.06195

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning

2026-04-03 · Cheng Gao, Cheng Huang, Kangyang Luo, Ziqing Qiao 외 arxiv

Enabling large language models (LLMs) to appropriately abstain from answering questions beyond their knowledge is crucial for mitigating hallucinations. While existing reinforcement learning methods foster autonomous abs…

Reinforcement Learning

Purging the Gray Zone: Latent-Geometric Denoising for Precise Knowledge Boundary Awareness

2026-04-15 · Hao An, Yibin Lou, Jiayi Guo, Yang Xu arxiv

Large language models (LLMs) often exhibit hallucinations due to their inability to accurately perceive their own knowledge boundaries. Existing abstention fine-tuning methods typically partition datasets directly based …

TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning

2025-09-30 · Zhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang 외 arxiv

While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside th…

General Reinforcement LearningQuestion Answering

Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

2026-05-30 · S M Tahmid Siddiqui, Akib Jawad Ononto, Anoop Singhal, Latifur Khan arxiv

Large language models (LLMs) have seen widespread adoption across various domains, yet their reliability is frequently undermined by hallucinations - responses that are plausible-sounding but factually incorrect. In high…

Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis

2025-12-16 · Richard Ackermann, Simeon Emanuilov arxiv

OpenAI has recently argued that hallucinations in large language models result primarily from misaligned evaluation incentives that reward confident guessing rather than epistemic humility. On this view, hallucination is…