paper-with-me

홈 › Papers

H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs

2025-12-01 · Cheng Gao, Huimin Chen, Chaojun Xiao, Zhiyi Chen, Zhiyuan Liu, Maosong Sun arxiv

Large language models (LLMs) frequently generate hallucinations -- plausible but factually incorrect outputs -- undermining their reliability. While prior work has examined hallucinations from macroscopic perspectives such as training data and objectives, the underlying neuron-level mechanisms remain largely unexplored. In this paper, we conduct a systematic investigation into hallucination-associated neurons (H-Neurons) in LLMs from three perspectives: identification, behavioral impact, and origins. Regarding their identification, we demonstrate that a remarkably sparse subset of neurons (less than $0.1\%$ of total neurons) can reliably predict hallucination occurrences, with strong generalization across diverse scenarios. In terms of behavioral impact, controlled interventions reveal that these neurons are causally linked to over-compliance behaviors. Concerning their origins, we trace these neurons back to the pre-trained base models and find that these neurons remain predictive for hallucination detection, indicating they emerge during pre-training. Our findings bridge macroscopic behavioral patterns with microscopic neural mechanisms, offering insights for developing more reliable LLMs.

📄 PDF Abstract BibTeX arXiv:2512.01797

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination

2026-06-30 · Vijay Vankadaru, Asha Matthews, Tanya Roosta, Peyman Passban arxiv

Hallucination remains one of the central obstacles to deploying medical LLMs. Yet, even when hallucination can be detected, it is still unclear whether the internal representations associated with it can be used for cont…

Finding Culture-Sensitive Neurons in Vision-Language Models

2025-10-28 · Xiutian Zhao, Rochelle Choenni, Rohit Saxena, Ivan Titov arxiv

Despite their impressive performance, vision-language models (VLMs) still struggle on culturally situated inputs. To understand how VLMs process culturally grounded information, we study the presence of culture-sensitive…

Visual Question Answering

MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models

2025-08-01 · Jiale Li, Mingrui Wu, Zixiang Jin, Hao Chen 외 arxiv

Despite growing interest in hallucination in Multimodal Large Language Models, existing studies primarily focus on single-image settings, leaving hallucination in multi-image scenarios largely unexplored. To address this…

Style-Specific Neurons for Steering LLMs in Text Style Transfer

2024-10-01 · Wen Lai, Viktor Hangya, Alexander Fraser

Text style transfer (TST) aims to modify the style of a text without altering its original meaning. Large language models (LLMs) demonstrate superior performance across multiple tasks, including TST. However, in zero-sho…

DiversityStyle TransferText Style Transfer

Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering

2026-01-31 · Guangtao Lyu, Xinyi Cheng, Qi Liu, Chenghao Xu 외 arxiv

LVLMs achieve remarkable multimodal understanding and generation but remain susceptible to hallucinations. Existing mitigation methods predominantly focus on output-level adjustments, leaving the internal mechanisms that…

Visual Grounding