paper-with-me

홈 › Papers

FACTUM: Mechanistic Detection of Citation Hallucination in Long-Form RAG

2026-01-09 · Maxime Dassen, Rebecca Kotula, Kenton Murray, Andrew Yates, Dawn Lawrie, Efsun Kayi, James Mayfield, Kevin Duh arxiv

Retrieval-Augmented Generation (RAG) models are critically undermined by citation hallucinations, a deceptive failure where a model cites a source that fails to support its claim. While existing work attributes hallucination to a simple over-reliance on parametric knowledge, we reframe this failure as an evolving, scale-dependent coordination failure between the Attention (reading) and Feed-Forward Network (recalling) pathways. We introduce FACTUM (Framework for Attesting Citation Trustworthiness via Underlying Mechanisms), a framework of four mechanistic scores: Contextual Alignment (CAS), Attention Sink Usage (BAS), Parametric Force (PFS), and Pathway Alignment (PAS). Our analysis reveals that correct citations are consistently marked by higher parametric force (PFS) and greater use of the attention sink (BAS) for information synthesis. Crucially, we find that "one-size-fits-all" theories are insufficient as the signature of correctness evolves with scale: while the 3B model relies on high pathway alignment (PAS), our best-performing 8B detector identifies a shift toward a specialized strategy where pathways provide distinct, orthogonal information. By capturing this complex interplay, FACTUM outperforms state-of-the-art baselines by up to 37.5% in AUC. Our results demonstrate that high parametric force is constructive when successfully coordinated with the Attention pathway, paving the way for more nuanced and reliable RAG systems.

📄 PDF Abstract BibTeX arXiv:2601.05866

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025

2026-02-05 · Samar Ansari arxiv

Large language models (LLMs) are increasingly used in academic writing workflows, yet they frequently hallucinate by generating citations to sources that do not exist. This study analyzes 100 AI-generated hallucinated ci…

Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective

2025-05-19 · Zhongxiang Sun, QiPeng Wang, Haoyu Wang, Xiao Zhang 외

Large Reasoning Models (LRMs) have shown impressive capabilities in multi-step reasoning tasks. However, alongside these successes, a more deceptive form of model error has emerged--Reasoning Hallucination--where logical…

Hallucination

InterpDetect: Interpretable Signals for Detecting Hallucinations in Retrieval-Augmented Generation

2025-10-24 · Likun Tan, Kuan-Wei Huang, Joy Shi, Kevin Wu arxiv

Retrieval-Augmented Generation (RAG) integrates external knowledge to mitigate hallucinations, yet models often generate outputs inconsistent with retrieved content. Accurate hallucination detection requires disentanglin…

Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations

2024-03-27 · Lei Yu, Meng Cao, Jackie Chi Kit Cheung, Yue Dong

State-of-the-art language models (LMs) sometimes generate non-factual hallucinations that misalign with world knowledge. To explore the mechanistic causes of these hallucinations, we create diagnostic datasets with subje…

AttributeDiagnosticHallucinationLanguage Modeling+3

CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text

2026-05-26 · Khashayar Khajavi, Shaghayegh Sadeghi, Rise Adhikari, Alexander Tessier arxiv

Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while containing corrupted metadata or pointing to papers that do not exist. We int…