paper-with-me

홈 › Papers

Markovian Circuit Tracing for Transformer State Dynamic

2026-05-20 · Abdullah X arxiv

Many sequence computations are easier to study as movement through internal states than as isolated local circuits. We introduce Markovian Circuit Tracing (MCT), a diagnostic pipeline for testing whether transformer activations contain coarse state-transition structure. The benchmark uses synthetic Hidden Markov Model (HMM) tasks where latent states, transition matrices, Bayesian belief vectors, Bayes-optimal predictions, and forced-state counterfactual targets are known exactly. Across six HMM families and three seeds per family, tiny causal transformers learn near-Bayes next-token predictors, with mean excess loss over Bayes of 0.0138. Residual activations contain partial Bayesian belief information in this controlled synthetic benchmark. State abstractions extracted from these activations recover coarse transition signal, strongest in persistent and lower-state regimes, and weaker in ambiguous-emission and six-state regimes. The clearest result comes from state forcing. Patching a recovered-state centroid reduces KL to the exact HMM counterfactual target from 0.1957 in the unpatched model to 0.0532 on average, beating wrong-state, mean-activation, random-activation, and shuffled-label controls. The contribution is a controlled benchmark and evaluation framework for transformer state-dynamics interpretability, with MCT as a simple reference pipeline

📄 PDF Abstract BibTeX arXiv:2605.20824

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing

2026-06-14 · Artyom Mazur, Nina Konovalova, Aibek Alanov arxiv

Mechanistic interpretability seeks to explain neural network behavior by decomposing model computations into interpretable features and circuits. While transcoder-based circuit tracing has recently enabled detailed causa…

Image Generation

Circuit Complexity of Hierarchical Knowledge Tracing and Implications for Log-Precision Transformers

2026-03-25 · Naiming Liu, Richard Baraniuk, Shashank Sonkar arxiv

Knowledge tracing models mastery over interconnected concepts, often organized by prerequisites. We analyze hierarchical prerequisite propagation through a circuit-complexity lens to clarify what is provable about transf…

Knowledge Tracing

Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing

2025-09-24 · Xinnan Dai, Chung-Hsiang Lo, Kai Guo, Shenglai Zeng 외 arxiv

Transformer-based LLMs demonstrate strong performance on graph reasoning tasks, yet their internal mechanisms remain underexplored. To uncover these reasoning process mechanisms in a fundamental and unified view, we set …

Anatomy of an Idiom: Tracing Non-Compositionality in Language Models

2025-11-20 · Andrew Gomes arxiv

We investigate the processing of idiomatic expressions in transformer-based language models using a novel set of techniques for circuit discovery and analysis. First discovering circuits via a modified path patching algo…

Computational Efficiency

Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking

2026-02-23 · Jingcheng Yang, Tianhu Xiong, Shengyi Qian, Klara Nahrstedt 외 arxiv

Vision-language models (VLMs) are powerful but remain opaque black boxes. We introduce the first framework for transparent circuit tracing in VLMs to systematically analyze multimodal reasoning. By utilizing transcoders,…

Mathematical ReasoningMultimodal Reasoning