paper-with-me

홈 › Papers

Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs

2025-12-13 · Ely Hahami, Ishaan Sinha, Lavik Jain, Josh Kaplan, Jon Hahami arxiv

Can large language models introspect, that is, accurately detect perturbations to their own internal states? We systematically investigate this question using activation steering in Meta-Llama-3.1-8B-Instruct. First, we show that the binary detection paradigm used in prior work conflates introspection with a methodological artifact: apparent detection accuracy is entirely explained by global logit shifts that bias models toward affirmative responses regardless of question content. However, on tasks requiring differential sensitivity, we find robust evidence for partial introspection: models localize which of 10 sentences received an injection at up to 88\% accuracy (vs.\ 10\% chance) and discriminate relative injection strengths at 83\% accuracy (vs.\ 50\% chance). These capabilities are confined to early-layer injections and collapse to chance thereafter -- a pattern we explain mechanistically through attention-based signal routing and residual stream recovery dynamics. Our findings demonstrate that LLMs can compute meaningful functions over perturbations to their internal states, establishing introspection as a real but layer-dependent phenomenon that merits further investigation. Our code is open-sourced here: https://github.com/elyhahami18/llama-introspection-new

📄 PDF Abstract BibTeX arXiv:2512.12411

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Metacognitive Prompting Improves Understanding in Large Language Models

2023-08-10 · Yuqing Wang, Yun Zhao

In Large Language Models (LLMs), there have been consistent advancements in task-specific performance, largely influenced by effective prompt design. Recent advancements in prompting have enhanced reasoning in logic-inte…

Natural Language Understanding

Agent Assessment of Others Through the Lens of Self

2023-12-18 · Jasmine A. Berry

The maturation of cognition, from introspection to understanding others, has long been a hallmark of human development. This position paper posits that for AI systems to truly emulate or approach human-like interactions,…

Computational EfficiencyPosition

AutoReply: Detecting Nonsense in Dialogue Introspectively with Discriminative Replies

2022-11-22 · Weiyan Shi, Emily Dinan, Adi Renduchintala, Daniel Fried 외

Existing approaches built separate classifiers to detect nonsense in dialogues. In this paper, we show that without external classifiers, dialogue models can detect errors in their own messages introspectively, by calcul…

H_\infty Almost Output and Regulated Output Synchronization of Heterogeneous Multi-agent Systems: A Scale-free Protocol Design

2021-04-16 · Donya Nojavanzadeh, Zhenwei Liu, Ali Saberi, Anton A. Stoorvogel

This paper studies scale-free protocol design for H_\infty almost output and regulated output synchronization of heterogeneous multi-agent systems with linear, right-invertible, and introspective agents in presence of ex…

Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

2024-08-04 · Fushuo Huo, Wenchao Xu, Zhong Zhang, Haozhao Wang 외

While Large Vision-Language Models (LVLMs) have rapidly advanced in recent years, the prevalent issue known as the `hallucination' problem has emerged as a significant bottleneck, hindering their real-world deployments. …

Hallucination