Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
Can large language models introspect, that is, accurately detect perturbations to their own internal states? We systematically investigate this question using activation steering in Meta-Llama-3.1-8B-Instruct. First, we show that the binary detection paradigm used in prior work conflates introspection with a methodological artifact: apparent detection accuracy is entirely explained by global logit shifts that bias models toward affirmative responses regardless of question content. However, on tasks requiring differential sensitivity, we find robust evidence for partial introspection: models localize which of 10 sentences received an injection at up to 88\% accuracy (vs.\ 10\% chance) and discriminate relative injection strengths at 83\% accuracy (vs.\ 50\% chance). These capabilities are confined to early-layer injections and collapse to chance thereafter -- a pattern we explain mechanistically through attention-based signal routing and residual stream recovery dynamics. Our findings demonstrate that LLMs can compute meaningful functions over perturbations to their internal states, establishing introspection as a real but layer-dependent phenomenon that merits further investigation. Our code is open-sourced here: https://github.com/elyhahami18/llama-introspection-new
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Metacognitive Prompting Improves Understanding in Large Language Models
In Large Language Models (LLMs), there have been consistent advancements in task-specific performance, largely influenced by effective prompt design. Recent advancements in prompting have enhanced reasoning in logic-inte…
Natural Language UnderstandingAgent Assessment of Others Through the Lens of Self
The maturation of cognition, from introspection to understanding others, has long been a hallmark of human development. This position paper posits that for AI systems to truly emulate or approach human-like interactions,…
Computational EfficiencyPositionAutoReply: Detecting Nonsense in Dialogue Introspectively with Discriminative Replies
Existing approaches built separate classifiers to detect nonsense in dialogues. In this paper, we show that without external classifiers, dialogue models can detect errors in their own messages introspectively, by calcul…
H_\infty Almost Output and Regulated Output Synchronization of Heterogeneous Multi-agent Systems: A Scale-free Protocol Design
This paper studies scale-free protocol design for H_\infty almost output and regulated output synchronization of heterogeneous multi-agent systems with linear, right-invertible, and introspective agents in presence of ex…
Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
While Large Vision-Language Models (LVLMs) have rapidly advanced in recent years, the prevalent issue known as the `hallucination' problem has emerged as a significant bottleneck, hindering their real-world deployments. …
Hallucination