paper-with-me

홈 › Papers

On the Faithfulness of Post-Hoc Concept Bottleneck Models

2026-06-29 · Laines Schmalwasser, Jan Blunk, Niklas Penzel, Julia Niebling, Joachim Denzler arxiv

Human decision-making interprets the world through high-level concepts, such as recognizing a bird by its belly color. To bridge the gap between opaque deep learning representations and human understanding, Post-Hoc Concept Bottleneck Models (post-hoc CBMs) project latent features onto interpretable concept spaces using auxiliary datasets or vision-language models. However, relying on target task accuracy as the primary measure of post-hoc CBM success obscures whether the learned concepts are semantically meaningful or merely predictive artifacts. For example, random concept projections can achieve competitive accuracy despite being semantically meaningless. In this work, we analyze the learned projections directly and identify two failure cases: First, for concept projections learned from auxiliary data, covariate shifts can lead to unfaithful concept representations for the target task. In particular, we provide an upper bound on the error introduced by this shift. Second, systematic label noise in surrogate concept labels generated by vision-language models leads to unfaithful projections. After formalizing these failure modes, we introduce novel metrics that decouple concept faithfulness from predictive accuracy. Our empirical results across real-world and synthetic benchmarks confirm that these metrics identify unfaithful behaviors that standard accuracy-based evaluation fails to detect.

📄 PDF Abstract BibTeX arXiv:2606.30498

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance

2024-07-18 · Divyansh Srivastava, Ge Yan, Tsui-Wei Weng

Concept Bottleneck Models (CBMs) provide interpretable prediction by introducing an intermediate Concept Bottleneck Layer (CBL), which encodes human-understandable concepts to explain models' decision. Recent works propo…

Avg

SL-CBM: Enhancing Concept Bottleneck Models with Semantic Locality for Better Interpretability

2026-01-19 · Hanwei Zhang, Luo Cheng, Rui Wen, Yang Zhang 외 arxiv

Explainable AI (XAI) is crucial for building transparent and trustworthy machine learning systems, especially in high-stakes domains. Concept Bottleneck Models (CBMs) have emerged as a promising ante-hoc approach that pr…

Do Concept Bottleneck Models Respect Localities?

2024-01-02 · Naveen Raman, Mateo Espinosa Zarlenga, Juyeon Heo, Mateja Jamnik

Concept-based methods explain model predictions using human-understandable concepts. These models require accurate concept predictors, yet the faithfulness of existing concept predictors to their underlying concepts is u…

Spatially Grounded Concept-Based Image Classification

2025-10-05 · Ran Eisenberg, Amit Rozner, Ethan Fetaya, Ofir Lindenbaum arxiv

Deep neural networks can achieve high accuracy while relying on evidence that is hard to inspect or misaligned with the intended task. Concept Bottleneck Models (CBMs) expose human-interpretable concepts, but most treat …

Image Classification

Towards Spatially-Aware and Optimally Faithful Concept-Based Explanations

2025-04-15 · Shubham Kumar, Dwip Dalal, Narendra Ahuja

Post-hoc, unsupervised concept-based explanation methods (U-CBEMs) are a promising tool for generating semantic explanations of the decision-making processes in deep neural networks, having applications in both model imp…

Decision Making