paper-with-me

홈 › Papers

Hallucination Detection via Activations of Open-Weight Proxy Analyzers

2026-05-08 · Akshita Singh, Prabesh Paudel, Siddhartha Roy arxiv

We introduce a proxy-analyzer framework for detecting hallucinations in large language models. Instead of looking inside the generating model, our system reads already-generated text through a small locally hosted open-weight model and spots hallucinations using the reader's own internal activations. This works just as well when the generator is a closed API like GPT-4 as when it is any open-weight model. We built eighteen features grounded in how transformers process text, covering residual stream norms, per-head source-document attention, entropy, MLP activations, logit-lens trajectories, and three new token-level grounding statistics. We trained a stacking ensemble on 72,135 samples from five hallucination datasets. We tested across seven analyzer architectures from 0.5 billion to 9 billion parameters: Qwen2.5 at 0.5B and 7B, Gemma-2 at 2B and 9B, Pythia at 1.4B, and LLaMA-3 at both 3B and 8B. Across all seven, we consistently beat ReDeEP's token-level AUC of 0.73 on RAGTruth by 7.4 to 10.3 percentage points. Qwen2.5-7B reached an F1 of 0.717, just above ReDeEP's 0.713, while Qwen2.5-0.5B hit 0.706. The most striking finding is how tightly all seven models cluster: AUC spans only 2.3 percentage points across an eighteen-fold difference in model size. Even more surprising, our 3B LLaMA outperforms our 8B LLaMA on RAGTruth, showing that bigger is not always better even within the same model family. Both RAGTruth and LLM-AggreFact include outputs from multiple LLM families, so our results are not skewed toward any particular generator.

📄 PDF Abstract BibTeX arXiv:2605.07209

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

D$^2$HScore: Reasoning-Aware Hallucination Detection via Semantic Breadth and Depth Analysis in LLMs

2025-09-15 · Yue Ding, Xiaofang Zhu, Tianze Xia, Junfei Wu 외 arxiv

Although large Language Models (LLMs) have achieved remarkable success, their practical application is often hindered by the generation of non-factual content, which is called "hallucination". Ensuring the reliability of…

On Early Detection of Hallucinations in Factual Question Answering

2023-12-19 · Ben Snyder, Marius Moisescu, Muhammad Bilal Zafar

While large language models (LLMs) have taken great strides towards helping humans with a plethora of tasks, hallucinations remain a major impediment towards gaining user trust. The fluency and coherence of model generat…

HallucinationOpen-Ended Question AnsweringQuestion Answering

Neural Message-Passing on Attention Graphs for Hallucination Detection

2025-09-29 · Fabrizio Frasca, Guy Bar-Shalom, Yftah Ziser, Haggai Maron arxiv

Large Language Models (LLMs) often generate incorrect or unsupported content, known as hallucinations. Existing detection methods rely on heuristics or simple models over isolated computational traces such as activations…

Graph Learning

EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs

2025-09-19 · Davide Ettori, Nastaran Darabi, Sina Tayebati, Ranganath Krishnan 외 arxiv

Large language models (LLMs) offer broad utility but remain prone to hallucination and out-of-distribution (OOD) errors. We propose EigenTrack, an interpretable real-time detector that uses the spectral geometry of hidde…

Out-of-Distribution Detection

Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination

2026-06-30 · Vijay Vankadaru, Asha Matthews, Tanya Roosta, Peyman Passban arxiv

Hallucination remains one of the central obstacles to deploying medical LLMs. Yet, even when hallucination can be detected, it is still unclear whether the internal representations associated with it can be used for cont…