paper-with-me

홈 › Papers

Shadows in the Attention: Contextual Perturbation and Representation Drift in the Dynamics of Hallucination in LLMs

2025-05-22 · Zeyu Wei, Shuo Wang, Xiaohui Rong, Xuemin Liu, He Li

Hallucinations -- plausible yet erroneous outputs -- remain a critical barrier to reliable deployment of large language models (LLMs). We present the first systematic study linking hallucination incidence to internal-state drift induced by incremental context injection. Using TruthfulQA, we construct two 16-round "titration" tracks per question: one appends relevant but partially flawed snippets, the other injects deliberately misleading content. Across six open-source LLMs, we track overt hallucination rates with a tri-perspective detector and covert dynamics via cosine, entropy, JS and Spearman drifts of hidden states and attention maps. Results reveal (1) monotonic growth of hallucination frequency and representation drift that plateaus after 5--7 rounds; (2) relevant context drives deeper semantic assimilation, producing high-confidence "self-consistent" hallucinations, whereas irrelevant context induces topic-drift errors anchored by attention re-routing; and (3) convergence of JS-Drift ($\sim0.69$) and Spearman-Drift ($\sim0$) marks an "attention-locking" threshold beyond which hallucinations solidify and become resistant to correction. Correlation analyses expose a seesaw between assimilation capacity and attention diffusion, clarifying size-dependent error modes. These findings supply empirical foundations for intrinsic hallucination prediction and context-aware mitigation mechanisms.

📄 PDF Abstract BibTeX arXiv:2505.16894

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationTruthfulQA

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Differential Analysis of Multispectral Images for Terrain Identification

2026-07-10 · Omar Kashmar, Hemendra Arya, Fulvio Mastrogiovanni arxiv

Reliable terrain understanding is a prerequisite for autonomous robot navigation. Yet, the widespread RGB-based perception can fail under low illumination, shadows, and material ambiguities. In this work we propose DRIFT…

Robot Navigation

Random Shadows and Highlights: A new data augmentation method for extreme lighting conditions

2021-01-13 · Osama Mazhar, Jens Kober

In this paper, we propose a new data augmentation method, Random Shadows and Highlights (RSH) to acquire robustness against lighting perturbations. Our method creates random shadows and highlights on images, thus challen…

Data Augmentation

HySAGE: A Hybrid Static and Adaptive Graph Embedding Network for Context-Drifting Recommendations

2022-08-20 · Sichun Luo, Xinyi Zhang, Yuanzhang Xiao, Linqi Song

The recent popularity of edge devices and Artificial Intelligent of Things (AIoT) has driven a new wave of contextual recommendations, such as location based Point of Interest (PoI) recommendations and computing resource…

Collaborative FilteringGraph Embedding

Deshadow-Anything: When Segment Anything Model Meets Zero-shot shadow removal

2023-09-21 · Xiao Feng Zhang, Tian Yi Song, Jia Wei Yao

Segment Anything (SAM), an advanced universal image segmentation model trained on an expansive visual dataset, has set a new benchmark in image segmentation and computer vision. However, it faced challenges when it came …

Image RestorationImage SegmentationImage Shadow RemovalSegmentation+2

Saliency Attention and Semantic Similarity-Driven Adversarial Perturbation

2024-06-18 · Hetvi Waghela, Jaydip Sen, Sneha Rakshit

In this paper, we introduce an enhanced textual adversarial attack method, known as Saliency Attention and Semantic Similarity driven adversarial Perturbation (SASSP). The proposed scheme is designed to improve the effec…

Adversarial AttackSemantic SimilaritySemantic Textual SimilaritySentence