paper-with-me

홈 › Papers

$α$-TCAV: A Unified Framework for Testing with Concept Activation Vectors

2026-05-15 · Ekkehard Schnoor, Jawher Said, Malik Tiomoko, Wojciech Samek, Alexander Jung arxiv

Concept Activation Vectors (CAVs) are a fundamental tool for concept-based explainability in deep learning, yet their practical utility is limited by statistical instability. We analyze the stochastic nature of CAVs and the Testing with CAVs (TCAV) method, deriving the distributions of major CAV classes including PatternCAV, FastCAV, and ridge regression-based CAVs. We then identify a fundamental flaw in the standard TCAV score: its reliance on a discontinuous indicator function induces non-decaying variance in critical regimes. To address this, we introduce $α$-TCAV, a generalized framework that replaces the indicator with a parameterized smooth function, yielding a unified probabilistic formulation that subsumes both TCAV and Multi-TCAV. We characterize the induced distributions of sensitivity scores and different TCAV variants, showing that established state-of-the-art choices lack theoretical justification. We provide principled guidance on tuning the parameter in $α$-TCAV -- either to imitate Multi-TCAV at substantially lower computational cost, or to obtain a calibrated Bayes-optimal probabilistic measure of a concept's influence. Finally, our analysis yields practical recommendations that challenge established routines: most notably, allocating the full sampling budget to a single CAV rather than splitting it across several.

📄 PDF Abstract BibTeX arXiv:2605.15688

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring Concept Contribution Spatially: Hidden Layer Interpretation with Spatial Activation Concept Vector

2022-05-21 · Andong Wang, Wei-Ning Lee

To interpret deep learning models, one mainstream is to explore the learned concepts by networks. Testing with Concept Activation Vector (TCAV) presents a powerful tool to quantify the contribution of query concepts (rep…

Visual-TCAV: Concept-based Attribution and Saliency Maps for Post-hoc Explainability in Image Classification

2024-11-08 · Antonio De Santis, Riccardo Campi, Matteo Bianchi, Marco Brambilla

Convolutional Neural Networks (CNNs) have seen significant performance improvements in recent years. However, due to their size and complexity, they function as black-boxes, leading to transparency concerns. State-of-the…

image-classificationImage Classification

E-TCAV: Formalizing Penultimate Proxies for Efficient Concept Based Interpretability

2026-05-11 · Hasib Aslam, Muhammad Ali Chattha, Muhammad Taha Mukhtar, Muhammad Imran Malik 외 arxiv

TCAV (Testing with Concept Activation Vectors) is an interpretability method that assesses the alignment between the internal representations of a trained neural network and human-understandable, high-level concepts. Tho…

TextCAVs: Debugging vision models using text

2024-08-16 · Angus Nicolson, Yarin Gal, J. Alison Noble

Concept-based interpretability methods are a popular form of explanation for deep learning models which provide explanations in the form of high-level human interpretable concepts. These methods typically find concept ac…

Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

2017-11-30 · ICML 2018 7 · Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai 외

The interpretation of deep learning models is a challenge due to their size, complexity, and often opaque internal state. In addition, many systems, such as image classifiers, operate on low-level features rather than hi…

General Classificationimage-classificationImage Classification