paper-with-me

홈 › Papers

Towards learning to explain with concept bottleneck models: mitigating information leakage

2022-11-07 · Joshua Lockhart, Nicolas Marchesotti, Daniele Magazzeni, Manuela Veloso

Concept bottleneck models perform classification by first predicting which of a list of human provided concepts are true about a datapoint. Then a downstream model uses these predicted concept labels to predict the target label. The predicted concepts act as a rationale for the target prediction. Model trust issues emerge in this paradigm when soft concept labels are used: it has previously been observed that extra information about the data distribution leaks into the concept predictions. In this work we show how Monte-Carlo Dropout can be used to attain soft concept predictions that do not contain leaked information.

📄 PDF Abstract BibTeX arXiv:2211.03656

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Hoeffding Concept Bottleneck Models with Applications to Overhead Images

2026-05-22 · Clément Bénard, Manon Arfib, Christophe Labreuche, Victor Quétu arxiv

Explainability of deep learning algorithms is critical for computer-vision applications with high-stake decisions. Concept bottleneck models (CBM) have recently shown promising performance to provide explainable and accu…

Object Detection

Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks

2026-06-17 · Ya Wang, Adrian Paschke arxiv

Concept Bottleneck Models (CBMs) enhance interpretability by projecting learned features into a human-understandable concept space. Recent approaches leverage vision-language models to generate concept embeddings, reduci…

Mitigating Bias in Concept Bottleneck Models for Fair and Interpretable Image Classification

2026-03-06 · Schrasing Tong, Antoine Salaun, Vincent Yuan, Annabel Adeyeri 외 arxiv

Ensuring fairness in image classification prevents models from perpetuating and amplifying bias. Concept bottleneck models (CBMs) map images to high-level, human-interpretable concepts before making predictions via a spa…

Image Classification

Promises and Pitfalls of Black-Box Concept Learning Models

2021-06-24 · Anita Mahinpei, Justin Clark, Isaac Lage, Finale Doshi-Velez 외

Machine learning models that incorporate concept learning as an intermediate step in their decision making process can match the performance of black-box predictive models while retaining the ability to explain outcomes …

Decision Making

Eliminating Information Leakage in Hard Concept Bottleneck Models with Supervised, Hierarchical Concept Learning

2024-02-03 · Ao Sun, Yuanyuan Yuan, Pingchuan Ma, Shuai Wang

Concept Bottleneck Models (CBMs) aim to deliver interpretable and interventionable predictions by bridging features and labels with human-understandable concepts. While recent CBMs show promising potential, they suffer f…