paper-with-me

홈 › Papers

Bucketing the Good Apples: A Method for Diagnosing and Improving Causal Abstraction

2026-05-04 · Li Puyin, Jiyuan Tan, Ahmad Jabbar, Thomas Icard, Atticus Geiger arxiv

We present a method for diagnosing interpretation in neural networks by identifying an input subspace where a proposed interpretation is highly faithful. Our method is particularly useful for causal-abstraction-style interpretability, where a high-level causal hypothesis is evaluated by interchange interventions. Rather than treating interchange intervention accuracy as a single global summary, we refine this framework by partitioning the input space into well-interpreted and under-interpreted regions according to pairwise interchange-intervention behavior. This turns causal abstraction from a purely global evaluation into a more diagnostic tool: it not only measures whether an interpretation works, but also reveals where it works, where it fails, and what distinguishes the two cases. This diagnostic view also provides practical heuristics for improving interpretations. By analyzing the structure of the well-interpreted and under-interpreted regions, we can identify missing distinctions in a high-level hypothesis, discover previously unmodeled intermediate variables, and combine complementary partial interpretations into a stronger one. We instantiate this idea as a simple four-step recipe and show that it yields informative error analyses across multiple causal abstraction settings. In a toy logic task, recursively applying the recipe recovers a high-level hypothesis from scratch. More broadly, our results suggest that partitioning the input space is a useful step toward more precise, constructive, and scalable mechanistic interpretability.

📄 PDF Abstract BibTeX arXiv:2605.02234

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Ordinal Bucketing for Game Trees using Dynamic Quantile Approximation

2019-05-31 · Tobias Joppen, Tilman Strübig, Johannes Fürnkranz

In this paper, we present a simple and cheap ordinal bucketing algorithm that approximately generates $q$-quantiles from an incremental data stream. The bucketing is done dynamically in the sense that the amount of bucke…

Learning Causal Abstractions of Linear Structural Causal Models

2024-06-01 · Riccardo Massidda, Sara Magliacane, Davide Bacciu

The need for modelling causal knowledge at different levels of granularity arises in several settings. Causal Abstraction provides a framework for formalizing this problem by relating two Structural Causal Models at diff…

Causal Discovery

Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection

2026-03-28 · Jinhu Fu, Yihang Lou, Qingyi Si, Shudong Zhang 외 arxiv

Large Vision-Language Models (LVLMs) have achieved impressive performance across multimodal understanding and reasoning tasks, yet their internal safety mechanisms remain opaque and poorly controlled. In this work, we pr…

Causal and Compositional Abstraction

2026-02-18 · Robin Lorenz, Sean Tull arxiv

Abstracting from a low level to a more explanatory high level of description, and ideally while preserving causal structure, is fundamental to scientific practice, to causal inference problems, and to robust, efficient a…

Causal Inference

Aligning Graphical and Functional Causal Abstractions

2024-12-22 · Willem Schooltink, Fabio Massimo Zennaro

Causal abstractions allow us to relate causal models on different levels of granularity. To ensure that the models agree on cause and effect, frameworks for causal abstractions define notions of consistency. Two distinct…