paper-with-me

홈 › Papers

Emergent Introspection in AI is Content-Agnostic

2026-03-05 · Harvey Lederman, Kyle Mahowald arxiv

Introspection is a foundational cognitive ability, but its mechanism is not well understood. Recent work has shown that AI models can introspect. We study the mechanism of this introspection. We first extensively replicate Lindsey (2025)'s thought injection detection paradigm in large open-source models. We show that introspection in these models is content-agnostic: models can detect that an anomaly occurred even when they cannot reliably identify its content. The models confabulate injected concepts that are high-frequency and concrete (e.g., "apple"). They also require fewer tokens to detect an injection than to guess the correct concept (with wrong guesses coming earlier). We argue that a content-agnostic introspective mechanism is consistent with leading theories in philosophy and psychology.

📄 PDF Abstract BibTeX arXiv:2603.05414

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language Models Fail to Introspect About Their Knowledge of Language

2025-03-10 · Siyuan Song, Jennifer Hu, Kyle Mahowald

There has been recent interest in whether large language models (LLMs) can introspect about their own internal states. Such abilities would make LLMs more interpretable, and also validate the use of standard introspectiv…

Sentence

Introspection Learning

2019-02-27 · Chris R. Serrano, Michael A. Warren

Traditional reinforcement learning agents learn from experience, past or present, gained through interaction with their environment. Our approach synthesizes experience, without requiring an agent to interact with their …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs

2025-12-13 · Ely Hahami, Ishaan Sinha, Lavik Jain, Josh Kaplan 외 arxiv

Can large language models introspect, that is, accurately detect perturbations to their own internal states? We systematically investigate this question using activation steering in Meta-Llama-3.1-8B-Instruct. First, we …

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention

2025-09-01 · Xiaomeng Hu, Fei Huang, Chenhan Yuan, Junyang Lin 외 arxiv

As large language models (LLMs) are increasingly deployed in real-world applications, ensuring the safety of their outputs during decoding has become a critical challenge. However, existing decoding-time interventions, s…

RIV: Recursive Introspection Mask Diffusion Vision Language Model

2025-09-28 · YuQian Li, Limeng Qiao, Lin Ma arxiv

Mask Diffusion-based Vision Language Models (MDVLMs) have achieved remarkable progress in multimodal understanding tasks. However, these models are unable to correct errors in generated tokens, meaning they lack self-cor…