paper-with-me

홈 › Papers

Causal Abstraction in Model Interpretability: A Compact Survey

2024-10-26 · Yihao Zhang

The pursuit of interpretable artificial intelligence has led to significant advancements in the development of methods that aim to explain the decision-making processes of complex models, such as deep learning systems. Among these methods, causal abstraction stands out as a theoretical framework that provides a principled approach to understanding and explaining the causal mechanisms underlying model behavior. This survey paper delves into the realm of causal abstraction, examining its theoretical foundations, practical applications, and implications for the field of model interpretability.

📄 PDF Abstract BibTeX arXiv:2410.20161

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingSurvey

Similar Papers 제목 키워드 기반

Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

2023-01-11 · Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary 외

Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level detai…

Explainable Artificial Intelligence (XAI)

The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?

2025-07-11 · Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel arxiv

The concept of causal abstraction got recently popularised to demystify the opaque decision-making processes of machine learning models; in short, a neural network can be abstracted as a higher-level algorithm if there e…

Interpreting Language Models Through Concept Descriptions: A Survey

2025-10-01 · Nils Feldhus, Laura Kopf arxiv

Understanding the decision-making processes of neural networks is a central goal of mechanistic interpretability. In the context of Large Language Models (LLMs), this involves uncovering the underlying mechanisms and ide…

Learning Causal Abstractions of Linear Structural Causal Models

2024-06-01 · Riccardo Massidda, Sara Magliacane, Davide Bacciu

The need for modelling causal knowledge at different levels of granularity arises in several settings. Causal Abstraction provides a framework for formalizing this problem by relating two Structural Causal Models at diff…

Causal Discovery

The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability

2024-08-02

Interpretability provides a toolset for understanding how and why neural networks behave in certain ways. However, there is little unity in the field: most studies employ ad-hoc evaluations and do not share theoretical f…