paper-with-me

홈 › Papers

The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?

2025-07-11 · Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel arxiv

The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level algorithm if there exists a function which allows us to map between them. Notably, most interpretability papers implement these maps as linear functions, motivated by the linear representation hypothesis: the idea that features are encoded linearly in a model's representations. However, this linearity constraint is not required by the definition of causal abstraction. In this work, we critically examine the concept of causal abstraction by considering arbitrarily powerful alignment maps. In particular, we prove that under reasonable assumptions, any neural network can be mapped to any algorithm, rendering this unrestricted notion of causal abstraction trivial and uninformative. We complement these theoretical findings with empirical evidence, demonstrating that it is possible to perfectly map models to algorithms even when these models are incapable of solving the actual task; e.g., on an experiment using randomly initialised language models, our alignment maps reach 100\% interchange-intervention accuracy on the indirect object identification task. This raises the non-linear representation dilemma: if we lift the linearity constraint imposed to alignment maps in causal abstraction analyses, we are left with no principled way to balance the inherent trade-off between these maps' complexity and accuracy. Together, these results suggest an answer to our title's question: causal abstraction is not enough for mechanistic interpretability, as it becomes vacuous without assumptions about how models encode information. Studying the connection between this information-encoding assumption and causal abstraction should lead to exciting future work.

📄 PDF Abstract BibTeX arXiv:2507.08802

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Causal Abstractions of Linear Structural Causal Models

2024-06-01 · Riccardo Massidda, Sara Magliacane, Davide Bacciu

The need for modelling causal knowledge at different levels of granularity arises in several settings. Causal Abstraction provides a framework for formalizing this problem by relating two Structural Causal Models at diff…

Causal Discovery

Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

2023-01-11 · Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary 외

Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level detai…

Explainable Artificial Intelligence (XAI)

CausalX: Causal Explanations and Block Multilinear Factor Analysis

2021-02-25 · M. Alex O. Vasilescu, Eric Kim, Xiao S. Zeng

By adhering to the dictum, "No causation without manipulation (treatment, intervention)", cause and effect data analysis represents changes in observed data in terms of changes in the causal factors. When causal factors …

Computational EfficiencycounterfactualObjectObject Recognition

Perturbation: A simple and efficient adversarial tracer for representation learning in language models

2026-03-25 · Joshua Rozner, Cory Shain arxiv

Linguistic representation learning in deep neural language models (LMs) has been studied for decades, for both practical and theoretical reasons. However, finding representations in LMs remains an unsolved problem, in pa…

Representation Learning

Causal Abstraction Inference under Lossy Representations

2025-09-25 · Kevin Xia, Elias Bareinboim arxiv

The study of causal abstractions bridges two integral components of human intelligence: the ability to determine cause and effect, and the ability to interpret complex patterns into abstract concepts. Formally, causal ab…