paper-with-me

홈 › Papers

Latent Introspection: Models Can Detect Prior Concept Injections

2026-02-23 · Theia Pearson-Vogel, Martin Vanek, Raymond Douglas, Jan Kulveit arxiv

We uncover a latent capacity for introspection in a Qwen 32B model, demonstrating that the model can detect when concepts have been injected into its earlier context and identify which concept was injected. While the model denies injection in sampled outputs, logit lens analysis reveals clear detection signals in the residual stream, which are attenuated in the final layers. Furthermore, prompting the model with accurate information about AI introspection mechanisms can dramatically strengthen this effect: the sensitivity to injection increases massively (0.3% -> 39.9%) with only a 0.6% increase in false positives. Also, mutual information between nine injected and recovered concepts rises from 0.61 bits to 1.05 bits, ruling out generic noise explanations. Our results demonstrate models can have a surprising capacity for introspection and steering awareness that is easy to overlook, with consequences for latent reasoning and safety.

📄 PDF Abstract BibTeX arXiv:2602.20031

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs

2025-12-13 · Ely Hahami, Ishaan Sinha, Lavik Jain, Josh Kaplan 외 arxiv

Can large language models introspect, that is, accurately detect perturbations to their own internal states? We systematically investigate this question using activation steering in Meta-Llama-3.1-8B-Instruct. First, we …

Emergent Introspection in AI is Content-Agnostic

2026-03-05 · Harvey Lederman, Kyle Mahowald arxiv

Introspection is a foundational cognitive ability, but its mechanism is not well understood. Recent work has shown that AI models can introspect. We study the mechanism of this introspection. We first extensively replica…

Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering

2026-01-08 · Shuliang Liu, Songbo Yang, Dong Fang, Sihang Jia 외 arxiv

Object hallucination critically undermines the reliability of Multimodal Large Language Models, often stemming from a fundamental failure in cognitive introspection, where models blindly trust linguistic priors over spec…

Me, Myself, and $π$ : Evaluating and Explaining LLM Introspection

2026-03-17 · Atharv Naphade, Samarth Bhargav, Sean Lim, Mcnair Shah arxiv

A hallmark of human intelligence is Introspection-the ability to assess and reason about one's own cognitive processes. Introspection has emerged as a promising but contested capability in large language models (LLMs). H…

Does It Make Sense to Speak of Introspection in Large Language Models?

2025-06-05 · Iulia M. Comsa, Murray Shanahan

Large language models (LLMs) exhibit compelling linguistic behaviour, and sometimes offer self-reports, that is to say statements about their own nature, inner workings, or behaviour. In humans, such reports are often at…

valid