paper-with-me

홈 › Papers

Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models

2024-03-11 · Weihang Su, Changyue Wang, Qingyao Ai, Yiran Hu, Zhijing Wu, Yujia Zhou, Yiqun Liu

Hallucinations in large language models (LLMs) refer to the phenomenon of LLMs producing responses that are coherent yet factually inaccurate. This issue undermines the effectiveness of LLMs in practical applications, necessitating research into detecting and mitigating hallucinations of LLMs. Previous studies have mainly concentrated on post-processing techniques for hallucination detection, which tend to be computationally intensive and limited in effectiveness due to their separation from the LLM's inference process. To overcome these limitations, we introduce MIND, an unsupervised training framework that leverages the internal states of LLMs for real-time hallucination detection without requiring manual annotations. Additionally, we present HELM, a new benchmark for evaluating hallucination detection across multiple LLMs, featuring diverse LLM outputs and the internal states of LLMs during their inference process. Our experiments demonstrate that MIND outperforms existing state-of-the-art methods in hallucination detection.

📄 PDF Abstract BibTeX arXiv:2403.06448

Code (2)

oneal2000/mind 공식 구현 pytorch
baylor-ai/halludetect

Tasks

Hallucination

Similar Papers 제목 키워드 기반

Unsupervised Hallucination Detection by Inspecting Reasoning Processes

2025-09-12 · Ponhvoan Srey, Xiaobao Wu, Anh Tuan Luu arxiv

Unsupervised hallucination detection aims to identify hallucinated content generated by large language models (LLMs) without relying on labeled data. While unsupervised methods have gained popularity by eliminating labor…

Internal Representations as Indicators of Hallucinations in Agent Tool Selection

2026-01-08 · Kait Healy, Bharathi Srinivasan, Visakh Madathil, Jing Wu arxiv

Large Language Models (LLMs) have shown remarkable capabilities in tool calling and tool usage, but suffer from hallucinations where they choose incorrect tools, provide malformed parameters and exhibit 'tool bypass' beh…

Prompt-Guided Internal States for Hallucination Detection of Large Language Models

2024-11-07 · Fujie Zhang, Peiqi Yu, Biao Yi, Baolei Zhang 외

Large Language Models (LLMs) have demonstrated remarkable capabilities across a variety of tasks in different domains. However, they sometimes generate responses that are logically coherent but factually incorrect or mis…

Domain GeneralizationHallucination

From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection

2026-06-22 · Jan Jasiński, Mateusz Barański, Julitta Bartolewska, Marcin Witkowski 외 arxiv

Hallucinations of ASR models - fluent transcriptions with no basis in audio - degrade system performance and pose risks in downstream applications. Robust detection of such errors remains a challenge. This paper studies …

Semantic Volume: Quantifying and Detecting both External and Internal Uncertainty in LLMs

2025-02-28 · Xiaomin Li, Zhou Yu, Ziji Zhang, Yingying Zhuang 외

Large language models (LLMs) have demonstrated remarkable performance across diverse tasks by encoding vast amounts of factual knowledge. However, they are still prone to hallucinations, generating incorrect or misleadin…

Hallucination