paper-with-me

Papers

Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims

2026-05-08 · Zezheng Lin, Fengming Liu arxiv

Mechanistic interpretability papers increasingly use causal vocabulary: circuits, mediators, causal abstraction, monosemanticity. Such claims require explicit identification assumptions. A purposive audit of 10 papers across four methodological strands finds no dedicated identification-assumptions section and a recurring pattern: validation metrics such as faithfulness, completeness, monosemanticity, alignment, or ablation effects are reported as causal support without stating the assumptions that make them identifying. A two-human-coder audit on $n=30$ reproduces the direction of the main finding: dedicated identification sections are absent, and validation-metric substitution is common, though exact Dim B/D counts are coding-rule sensitive. The paper proposes a disclosure norm: state whether the claim is causal, name the identification strategy, enumerate assumptions, stress at least one, and explain how conclusions shift if assumptions fail. Validation is not identification.

📄 PDF Abstract BibTeX arXiv:2605.08012

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Mechanistic to Compositional Interpretability

2026-05-09 · Ward Gauderis, Thomas Dooms, Steven T. Homer, Kola Ayonrinde 외 arxiv

Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable components. Without a formal framework, however, mechanistic explanatio…

Investigating the Indirect Object Identification circuit in Mamba

2024-07-19 · Danielle Ensign, Adrià Garriga-Alonso

How well will current interpretability techniques generalize to future models? A relevant case study is Mamba, a recent recurrent architecture with scaling comparable to Transformers. We adapt pre-Mamba techniques to Mam…

MambaObjectPosition

Open Problems in Mechanistic Interpretability

2025-01-27 · Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey 외

Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goals. Progress in this field thus promises…

Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition

2025-01-24 · Dan Braun, Lucius Bushnaq, Stefan Heimersheim, Jake Mendel 외

Mechanistic interpretability aims to understand the internal mechanisms learned by neural networks. Despite recent progress toward this goal, it remains unclear how best to decompose neural network parameters into mechan…

Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability

2025-08-20 · Ashwath Vaithinathan Aravindan, Abha Jha, Mihir Kulkarni arxiv

Vision-Language Models (VLMs) have shown remarkable performance in integrating visual and textual information for tasks such as image captioning and visual question answering. However, these models struggle with composit…

Visual Question AnsweringImage Captioning