paper-with-me

홈 › Papers

Mechanisms of Prompt-Induced Hallucination in Vision-Language Models

2026-01-08 · William Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov, Ritambhara Singh, Carsten Eickhoff, Kyle Mahowald arxiv

Large vision-language models (VLMs) are highly capable, yet often hallucinate by favoring textual prompts over visual evidence. We study this failure mode in a controlled object-counting setting, where the prompt overstates the number of objects in the image (e.g., asking a model to describe four waterlilies when only three are present). At low object counts, models often correct the overestimation, but as the number of objects increases, they increasingly conform to the prompt regardless of the discrepancy. Through mechanistic analysis of three VLMs, we identify a small set of attention heads whose ablation substantially reduces prompt-induced hallucinations (PIH) by at least 40% without additional training. Across models, PIH-heads mediate prompt copying in model-specific ways. We characterize these differences and show that PIH ablation increases correction toward visual evidence. Our findings offer insights into the internal mechanisms driving prompt-induced hallucinations, revealing model-specific differences in how these behaviors are implemented.

📄 PDF Abstract BibTeX arXiv:2601.05201

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs

2026-04-23 · Pegah Khayatan, Jayneel Parekh, Arnaud Dapogny, Mustafa Shukor 외 arxiv

Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the visual input. Prior work has attributed h…

The Way We Prompt: Conceptual Blending, Neural Dynamics, and Prompt-Induced Transitions in LLMs

2025-05-16 · Makoto Sato

Large language models (LLMs), inspired by neuroscience, exhibit behaviors that often evoke a sense of personality and intelligence-yet the mechanisms behind these effects remain elusive. Here, we operationalize Conceptua…

Prompt Engineering

Mitigating Prompt-Induced Hallucinations in Large Language Models via Structured Reasoning

2026-01-06 · Jinbo Hao, Kai Yang, Qingzhen Su, Yang Chen 외 arxiv

To address hallucination issues in large language models (LLMs), this paper proposes a method for mitigating prompt-induced hallucinations. Building on a knowledge distillation chain-style model, we introduce a code modu…

Knowledge Distillation

CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement

2025-10-14 · Jung-Woo Shim, Yeong-Joon Ju, Ji-Hoon Park, Seong-Whan Lee arxiv

Recent advancements in large language models (LLMs) highlight their fluency in generating responses to diverse prompts. However, these models sometimes generate plausible yet incorrect ``hallucinated" facts, undermining …

Neural Uncertainty Principle: A Unified View of Adversarial Fragility and LLM Hallucination

2026-03-20 · Dong-Xiao Zhang, Hu Lou, Jun-Jie Zhang, Jun Zhu 외 arxiv

Adversarial vulnerability in vision and hallucination in large language models are conventionally viewed as separate problems, each addressed with modality-specific patches. This study first reveals that they share a com…