paper-with-me

홈 › Papers

Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution

2025-09-19 · Zuzanna Dubanowska, Maciej Żelaszczyk, Michał Brzozowski, Paolo Mandica, Michał Karpowicz arxiv

We critically assess the efficacy of the current SOTA in hallucination detection and find that its performance on the RAGTruth dataset is largely driven by a spurious correlation with data. Controlling for this effect, state-of-the-art performs no better than supervised linear probes, while requiring extensive hyperparameter tuning across datasets. Out-of-distribution generalization is currently out of reach, with all of the analyzed methods performing close to random. We propose a set of guidelines for hallucination detection and its evaluation.

📄 PDF Abstract BibTeX arXiv:2509.19372

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Do LLM hallucination detectors suffer from low-resource effect?

2026-01-23 · Debtanu Datta, Mohan Kishore Chilukuri, Yash Kumar, Saptarshi Ghosh 외 arxiv

LLMs, while outperforming humans in a wide range of tasks, can still fail in unanticipated ways. We focus on two pervasive failure modes: (i) hallucinations, where models produce incorrect information about the world, an…

Robust Hallucination Detection in LLMs via Adaptive Token Selection

2025-04-10 · Mengjia Niu, Hamed Haddadi, Guansong Pang

Hallucinations in large language models (LLMs) pose significant safety concerns that impede their broader deployment. Recent research in hallucination detection has demonstrated that LLMs' internal representations contai…

Hallucination

LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

2024-10-03 · Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart 외

Large language models (LLMs) often produce errors, including factual inaccuracies, biases, and reasoning failures, collectively referred to as "hallucinations". Recent studies have demonstrated that LLMs' internal states…

Prompt-Guided Internal States for Hallucination Detection of Large Language Models

2024-11-07 · Fujie Zhang, Peiqi Yu, Biao Yi, Baolei Zhang 외

Large Language Models (LLMs) have demonstrated remarkable capabilities across a variety of tasks in different domains. However, they sometimes generate responses that are logically coherent but factually incorrect or mis…

Domain GeneralizationHallucination

EigenTrack: Spectral Activation Feature Tracking for Hallucination and Out-of-Distribution Detection in LLMs and VLMs

2025-09-19 · Davide Ettori, Nastaran Darabi, Sina Tayebati, Ranganath Krishnan 외 arxiv

Large language models (LLMs) offer broad utility but remain prone to hallucination and out-of-distribution (OOD) errors. We propose EigenTrack, an interpretable real-time detector that uses the spectral geometry of hidde…

Out-of-Distribution Detection