paper-with-me

Papers

Mechanisms of Introspective Awareness

2026-03-22 · Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey arxiv

Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introspective awareness." We investigate the mechanisms underlying this capability in open-weights models. First, we find that it is behaviorally robust: models detect injected steering vectors at moderate rates with 0% false positives across diverse prompts and dialogue formats. Notably, this capability emerges specifically from post-training; we show that preference optimization algorithms like DPO can elicit it, but standard supervised finetuning does not. We provide evidence that detection cannot be explained by simple linear association between certain steering vectors and directions promoting affirmative responses. We trace the detection mechanism to a two-stage circuit in which "evidence carrier" features in early post-injection layers detect perturbations monotonically along diverse directions, suppressing downstream "gate" features that implement a default negative response. This circuit is absent in base models and robust to refusal ablation. Identification of injected concepts relies on largely distinct later-layer mechanisms that only weakly overlap with those involved in detection. Finally, we show that introspective capability is substantially underelicited: ablating refusal directions improves detection by +53%, and a trained bias vector improves it by +75% on held-out concepts, both without meaningfully increasing false positives. Our results suggest that this introspective awareness of injected concepts is robust and mechanistically nontrivial, and could be substantially amplified in future models. Code: https://github.com/safety-research/introspection-mechanisms.

📄 PDF Abstract BibTeX arXiv:2603.21396

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploration Through Introspection: A Self-Aware Reward Model

2026-01-06 · Michael Petrowski, Milica Gašić arxiv

Understanding how artificial agents model internal mental states is central to advancing Theory of Mind in AI. Evidence points to a unified system for self- and other-awareness. We explore this self-awareness by having r…

Reinforcement Learning

Emergent Introspective Awareness in Large Language Models

2026-01-05 · Jack Lindsey arxiv

We investigate whether large language models can introspect on their internal states. It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulatio…

Introspective Deep Metric Learning

2023-09-11 · Chengkun Wang, Wenzhao Zheng, Zheng Zhu, Jie zhou 외

This paper proposes an introspective deep metric learning (IDML) framework for uncertainty-aware comparisons of images. Conventional deep metric learning methods focus on learning a discriminative embedding to describe t…

Image RetrievalMetric Learning

Introspective Perception: Learning to Predict Failures in Vision Systems

2016-07-28 · Shreyansh Daftry, Sam Zeng, J. Andrew Bagnell, Martial Hebert

As robots aspire for long-term autonomous operations in complex dynamic environments, the ability to reliably take mission-critical decisions in ambiguous situations becomes critical. This motivates the need to build sys…

Beyond the Wavefunction: Qualia Abstraction Language Mechanics and the Grammar of Awareness

2025-08-03 · Mikołaj Sienicki, Krzysztof Sienicki arxiv

We propose a formal reconstruction of quantum mechanics grounded not in external mathematical abstractions, but in the structured dynamics of subjective experience. The Qualia Abstraction Language (QAL) models physical s…