Describe-and-Dissect: Interpreting Neurons in Vision Networks with Language Models
In this paper, we propose Describe-and-Dissect (DnD), a novel method to describe the roles of hidden neurons in vision networks. DnD utilizes recent advancements in multimodal deep learning to produce complex natural language descriptions, without the need for labeled training data or a predefined set of concepts to choose from. Additionally, DnD is training-free, meaning we don't train any new models and can easily leverage more capable general purpose models in the future. We have conducted extensive qualitative and quantitative analysis to show that DnD outperforms prior work by providing higher quality neuron descriptions. Specifically, our method on average provides the highest quality labels and is more than 2 times as likely to be selected as the best explanation for a neuron than the best baseline.
Code (0)
등록된 구현이 없습니다.
Tasks
Multimodal Deep LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks
In this paper, we propose CLIP-Dissect, a new technique to automatically describe the function of individual hidden neurons inside vision networks. CLIP-Dissect leverages recent advances in multimodal vision/language mod…
AND: Audio Network Dissection for Interpreting Deep Acoustic Models
Neuron-level interpretations aim to explain network behaviors and properties by investigating neurons responsive to specific perceptual or structural input patterns. Although there is emerging work in the vision and lang…
Machine UnlearningSelf-Supervised LearningMammo-CLIP Dissect: A Framework for Analysing Mammography Concepts in Vision-Language Models
Understanding what deep learning (DL) models learn is essential for the safe deployment of artificial intelligence (AI) in clinical settings. While previous work has focused on pixel-based explainability methods, less at…
DISCOVER: Making Vision Networks Interpretable via Competition and Dissection
Modern deep networks are highly complex and their inferential outcome very hard to interpret. This is a serious obstacle to their transparent deployment in safety-critical or bias-aware applications. This work contribute…
Interpreting Neural Networks through the Polytope Lens
Mechanistic interpretability aims to explain what a neural network has learned at a nuts-and-bolts level. What are the fundamental primitives of neural network representations? Previous mechanistic descriptions have used…