paper-with-me

홈 › Papers

HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals

2026-08-17 · Zhihao Guo, Zonghan Wu, Huan Huo, DaYong Ye, Junwei Zhang, Weiran Yao, Zhiwei Liu, Qingsong Wen, Yilei Shao arxiv

Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments. These models nonetheless carry linearly separable truthfulness signals in their internal representations. Existing white-box detectors, however, collapse this evidence to isolated components or a single depth, discarding discriminative information distributed across the full forward pass. We introduce HalluTracer, a detection framework that reads and aggregates truthfulness evidence across every layer of the forward pass before the model emits any answer token. A geometric analysis reveals that the per-layer signals are weakly correlated, so that simple depth averaging suppresses layer-specific noise and captures nearly all linearly accessible information. Across six open-source language models and five hallucination benchmarks, HalluTracer consistently outperforms matched white-box baselines, with gains ranging from one to fourteen points. Collectively, our work recasts hallucination detection from a layer-selection problem into a depth-aggregation problem governed by the geometric sparsity of the truthfulness signal.

📄 PDF Abstract BibTeX arXiv:2608.16353

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TruthPrInt: Mitigating LVLM Object Hallucination Via Latent Truthful-Guided Pre-Intervention

2025-03-13 · Jinhao Duan, Fei Kong, Hao Cheng, James Diffenderfer 외

Object Hallucination (OH) has been acknowledged as one of the major trustworthy challenges in Large Vision-Language Models (LVLMs). Recent advancements in Large Language Models (LLMs) indicate that internal states, such …

HallucinationObject HallucinationSpecificity

The Geometry of Truth: Layer-wise Semantic Dynamics for Hallucination Detection in Large Language Models

2025-10-06 · Amir Hameed Mir arxiv

Large Language Models (LLMs) often produce fluent yet factually incorrect statements-a phenomenon known as hallucination-posing serious risks in high-stakes domains. We present Layer-wise Semantic Dynamics (LSD), a geome…

Contrastive Learning

Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective

2025-05-19 · Zhongxiang Sun, QiPeng Wang, Haoyu Wang, Xiao Zhang 외

Large Reasoning Models (LRMs) have shown impressive capabilities in multi-step reasoning tasks. However, alongside these successes, a more deceptive form of model error has emerged--Reasoning Hallucination--where logical…

Hallucination

FactCHD: Benchmarking Fact-Conflicting Hallucination Detection

2023-10-18 · Xiang Chen, Duanzheng Song, Honghao Gui, Chenxi Wang 외

Despite their impressive generative capabilities, LLMs are hindered by fact-conflicting hallucinations in real-world applications. The accurate identification of hallucinations in texts generated by LLMs, especially in c…

BenchmarkingHallucination

OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution

2026-09-06 · Shubhashis Roy Dipta, Sourajit Saha, Shaswati Saha, Nobin Sarwar hf

Recursive Super-Resolution (SR) extends fixed-scale SR to extreme magnification by repeatedly feeding predictions back into the same model, analogous to zooming an image repeatedly. However, ground truth availability at …