paper-with-me

Papers

Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation

2026-04-16 · G. Aytug Akarlar arxiv

We present causal evidence that hallucination in autoregressive language models is an early trajectory commitment governed by asymmetric attractor dynamics. Using same-prompt bifurcation, in which we repeatedly sample identical inputs to observe spontaneous divergence, we isolate trajectory dynamics from prompt-level confounds. On Qwen2.5-1.5B across 61 prompts spanning six categories, 27 prompts (44.3%) bifurcate with factual and hallucinated trajectories diverging at the first generated token (KL = 0 at step 0, KL > 1.0 at step 1). Activation patching across 28 layers reveals a pronounced causal asymmetry: injecting a hallucinated activation into a correct trajectory corrupts output in 87.5% of trials (layer 20), while the reverse recovers only 33.3% (layer 24); both exceed the 10.4% baseline (p = 0.025) and 12.5% random-patch control. Window patching shows correction requires sustained multi-step intervention, whereas corruption needs only a single perturbation. Probing the prompt encoding itself, step-0 residual states predict per-prompt hallucination rate at Pearson r = 0.776 at layer 15 (p < 0.001 against a 1000-permutation null); unsupervised clustering identifies five regime-like groups (eta^2 = 0.55) whose saddle-adjacent cluster concentrates 12 of the 13 bifurcating false-premise prompts, indicating that the basin structure is organized around regime commitments fixed at prompt encoding. These findings characterize hallucination as a locally stable attractor basin: entry is probabilistic and rapid, exit demands coordinated intervention across layers and steps, and the relevant basins are selected by clusterable regimes already discernible at step 0.

📄 PDF Abstract BibTeX arXiv:2604.15400

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OSCAR: Orchestrated Self-verification and Cross-path Refinement

2026-04-02 · Yash Shah, Abhijit Chakraborty, Naresh Kumar Devulapally, Vishnu Lokhande 외 arxiv

Diffusion language models (DLMs) expose their denoising trajectories, offering a natural handle for inference-time control; accordingly, an ideal hallucination mitigation framework should intervene during generation usin…

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination

2026-05-19 · Jinrui Jiang, Zhangtai Wu, Zhen Wu, Xinyu Dai arxiv

Modality-conflict hallucination occurs when multimodal large language models (MLLMs) prioritize erroneous textual premises over contradictory visual evidence. To understand why visual evidence fails to prevail during gen…

Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations

2026-08-01 · Xinshun Feng, Ziqi Miao, Lijun Li, Jing Shao arxiv

Large language model (LLM) agents are increasingly deployed in scientific research, where reliability is critical and the underlying knowledge is densely interconnected. In such settings, hallucinations are particularly …

Code Generation

Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning

2026-06-08 · Haoran Xu, Hongyu Wang, Yifei Gao, Jiaze Li 외 arxiv

Visual reasoning requires integrating evidence distributed across regions, attributes, and relations, making single-chain reasoning prone to early perceptual commitment and hallucination. We propose Visual Para-Thinker++…

Visual Reasoning

HART: Data-Driven Hallucination Attribution and Evidence-Based Tracing for Large Language Models

2026-03-06 · Shize Liang, Hongzhi Wang arxiv

Large language models (LLMs) have demonstrated remarkable performance in text generation and knowledge-intensive question answering. Nevertheless, they are prone to producing hallucinated content, which severely undermin…

Semantic SimilarityQuestion AnsweringText Generation