paper-with-me

홈 › Papers

Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective

2025-05-19 · Zhongxiang Sun, QiPeng Wang, Haoyu Wang, Xiao Zhang, Jun Xu

Large Reasoning Models (LRMs) have shown impressive capabilities in multi-step reasoning tasks. However, alongside these successes, a more deceptive form of model error has emerged--Reasoning Hallucination--where logically coherent but factually incorrect reasoning traces lead to persuasive yet faulty conclusions. Unlike traditional hallucinations, these errors are embedded within structured reasoning, making them more difficult to detect and potentially more harmful. In this work, we investigate reasoning hallucinations from a mechanistic perspective. We propose the Reasoning Score, which quantifies the depth of reasoning by measuring the divergence between logits obtained from projecting late layers of LRMs to the vocabulary space, effectively distinguishing shallow pattern-matching from genuine deep reasoning. Using this score, we conduct an in-depth analysis on the ReTruthQA dataset and identify two key reasoning hallucination patterns: early-stage fluctuation in reasoning depth and incorrect backtracking to flawed prior steps. These insights motivate our Reasoning Hallucination Detection (RHD) framework, which achieves state-of-the-art performance across multiple domains. To mitigate reasoning hallucinations, we further introduce GRPO-R, an enhanced reinforcement learning algorithm that incorporates step-level deep reasoning rewards via potential-based shaping. Our theoretical analysis establishes stronger generalization guarantees, and experiments demonstrate improved reasoning quality and reduced hallucination rates.

📄 PDF Abstract BibTeX arXiv:2505.12886

Code (0)

등록된 구현이 없습니다.

Tasks

Hallucination

Similar Papers 제목 키워드 기반

Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations

2024-03-27 · Lei Yu, Meng Cao, Jackie Chi Kit Cheung, Yue Dong

State-of-the-art language models (LMs) sometimes generate non-factual hallucinations that misalign with world knowledge. To explore the mechanistic causes of these hallucinations, we create diagnostic datasets with subje…

AttributeDiagnosticHallucinationLanguage Modeling+3

Dissecting the Ledger: Locating and Suppressing "Liar Circuits" in Financial Large Language Models

2025-11-24 · Soham Mirajkar arxiv

Large Language Models (LLMs) are increasingly deployed in high-stakes financial domains, yet they suffer from specific, reproducible hallucinations when performing arithmetic operations. Current mitigation strategies oft…

Arithmetic Reasoning

Context-Aware Decoding for Faithful Vision-Language Generation

2026-01-09 · Mehrdad Fazli, Bowen Wei, Ziwei Zhu arxiv

Hallucinations, generating responses inconsistent with the visual input, remain a critical limitation of large vision-language models (LVLMs), especially in open-ended tasks such as image captioning and visual reasoning.…

Visual ReasoningImage Captioning

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination

2025-10-27 · Chenlong Yin, Zeyang Sha, Shiwen Cui, Changhua Meng 외 arxiv

Enhancing the reasoning capabilities of Large Language Models (LLMs) is a key strategy for building Agents that "think then act." However, recent observations, like OpenAI's o3, suggest a paradox: stronger reasoning ofte…

Prompt Engineering

Hallucination Detection and Mitigation in Large Language Models

2026-01-14 · Ahmad Pesaranghader, Erin Li arxiv

Large Language Models (LLMs) and Large Reasoning Models (LRMs) offer transformative potential for high-stakes domains like finance and law, but their tendency to hallucinate, generating factually incorrect or unsupported…