paper-with-me

홈 › Papers

Prototype-Grounded Concept Models for Verifiable Concept Alignment

2026-04-17 · Stefano Colamonaco, David Debot, Pietro Barbiero, Giuseppe Marra arxiv

Concept Bottleneck Models (CBMs) aim to improve interpretability in Deep Learning by structuring predictions through human-understandable concepts, but they provide no way to verify whether learned concepts align with the human's intended meaning, hurting interpretability. We introduce Prototype-Grounded Concept Models (PGCMs), which ground concepts in learned visual prototypes: image parts that serve as explicit evidence for the concepts. This grounding enables direct inspection of concept semantics and supports targeted human intervention at the prototype level to correct misalignments. Empirically, PGCMs achieve similar predictive performance as state-of-the-art CBMs while substantially improving transparency, interpretability, and intervenability.

📄 PDF Abstract BibTeX arXiv:2604.16076

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interactive Disentanglement: Learning Concepts by Interacting with their Prototype Representations

2021-12-04 · CVPR 2022 1 · Wolfgang Stammer, Marius Memmel, Patrick Schramowski, Kristian Kersting

Learning visual concepts from raw images without strong supervision is a challenging task. In this work, we show the advantages of prototype representations for understanding and revising the latent space of neural conce…

Disentanglement

Beyond Patches: Mining Interpretable Part-Prototypes for Explainable AI

2025-04-16 · Mahdi Alehdaghi, Rajarshi Bhattacharya, Pourya Shamsolmoali, Rafael M. O. Cruz 외

Deep learning has provided considerable advancements for multimedia systems, yet the interpretability of deep models remains a challenge. State-of-the-art post-hoc explainability methods, such as GradCAM, provide visual …

Unsupervised Part Discovery

Using Grounded Word Representations to Study Theories of Lexical Concepts

2019-06-01 · WS 2019 6 · Dylan Ebert, Ellie Pavlick

The fields of cognitive science and philosophy have proposed many different theories for how humans represent {``}concepts{''}. Multiple such theories are compatible with state-of-the-art NLP methods, and could in princi…

Philosophy

MCPNet: An Interpretable Classifier via Multi-Level Concept Prototypes

2024-04-13 · CVPR 2024 1 · Bor-Shiun Wang, Chien-Yi Wang, Wei-Chen Chiu

Recent advancements in post-hoc and inherently interpretable methods have markedly enhanced the explanations of black box classifier models. These methods operate either through post-analysis or by integrating concept le…

ClassificationDecision MakingExplainable Artificial Intelligence (XAI)

AI Integrity: A New Paradigm for Verifiable AI Governance

2026-04-13 · Seulki Lee arxiv

AI systems increasingly shape high-stakes decisions in healthcare, law, defense, and education, yet existing governance paradigms -- AI Ethics, AI Safety, and AI Alignment -- share a common limitation: they evaluate outc…