paper-with-me

홈 › Papers

Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence

2022-02-07 · Frederik Pahde, Maximilian Dreyer, Leander Weber, Moritz Weckbecker, Christopher J. Anders, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin

With a growing interest in understanding neural network prediction strategies, Concept Activation Vectors (CAVs) have emerged as a popular tool for modeling human-understandable concepts in the latent space. Commonly, CAVs are computed by leveraging linear classifiers optimizing the separability of latent representations of samples with and without a given concept. However, in this paper we show that such a separability-oriented computation leads to solutions, which may diverge from the actual goal of precisely modeling the concept direction. This discrepancy can be attributed to the significant influence of distractor directions, i.e., signals unrelated to the concept, which are picked up by filters (i.e., weights) of linear models to optimize class-separability. To address this, we introduce pattern-based CAVs, solely focussing on concept signals, thereby providing more accurate concept directions. We evaluate various CAV methods in terms of their alignment with the true concept direction and their impact on CAV applications, including concept sensitivity testing and model correction for shortcut behavior caused by data artifacts. We demonstrate the benefits of pattern-based CAVs using the Pediatric Bone Age, ISIC2019, and FunnyBirds datasets with VGG, ResNet, ReXNet, EfficientNet, and Vision Transformer as model architectures.

📄 PDF Abstract BibTeX arXiv:2202.03482

Code (0)

등록된 구현이 없습니다.

Tasks

TAG

Similar Papers 제목 키워드 기반

Robust Semantic Interpretability: Revisiting Concept Activation Vectors

2021-04-06 · Jacob Pfau, Albert T. Young, Jerome Wei, Maria L. Wei 외

Interpretability methods for image classification assess model trustworthiness by attempting to expose whether the model is systematically biased or attending to the same cues as a human would. Saliency methods for featu…

Benchmarkingcounterfactualimage-classificationImage Classification

Concept activation vectors: a unifying view and adversarial attacks

2025-09-26 · Ekkehard Schnoor, Malik Tiomoko, Jawher Said, Alex Jung 외 arxiv

Concept Activation Vectors (CAVs) are a tool from explainable AI, offering a promising approach for understanding how human-understandable concepts are encoded in a model's latent spaces. They are computed from hidden-la…

Adversarial Attack

Concept Boundary Vectors

2024-12-20 · Thomas Walker

Machine learning models are trained with relatively simple objectives, such as next token prediction. However, on deployment, they appear to capture a more fundamental representation of their input data. It is of interes…

Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors

2026-08-06 · Arya Labroo, Mengjie Qian, Kate Knill arxiv

Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency r…

Concept Activation Vectors for Generating User-Defined 3D Shapes

2022-04-29 · Stefan Druc, Aditya Balu, Peter Wooldridge, Adarsh Krishnamurthy 외

We explore the interpretability of 3D geometric deep learning models in the context of Computer-Aided Design (CAD). The field of parametric CAD can be limited by the difficulty of expressing high-level design concepts in…