paper-with-me

홈 › Papers

Interpreting the Second-Order Effects of Neurons in CLIP

2024-06-06 · Yossi Gandelsman, Alexei A. Efros, Jacob Steinhardt

We interpret the function of individual neurons in CLIP by automatically describing them using text. Analyzing the direct effects (i.e. the flow from a neuron through the residual stream to the output) or the indirect effects (overall contribution) fails to capture the neurons' function in CLIP. Therefore, we present the "second-order lens", analyzing the effect flowing from a neuron through the later attention heads, directly to the output. We find that these effects are highly selective: for each neuron, the effect is significant for <2% of the images. Moreover, each effect can be approximated by a single direction in the text-image space of CLIP. We describe neurons by decomposing these directions into sparse sets of text representations. The sets reveal polysemantic behavior - each neuron corresponds to multiple, often unrelated, concepts (e.g. ships and cars). Exploiting this neuron polysemy, we mass-produce "semantic" adversarial examples by generating images with concepts spuriously correlated to the incorrect class. Additionally, we use the second-order effects for zero-shot segmentation and attribute discovery in images. Our results indicate that a scalable understanding of neurons can be used for model deception and for introducing new model capabilities.

📄 PDF Abstract BibTeX arXiv:2406.04341

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeZero Shot Segmentation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Interpreting ResNet-based CLIP via Neuron-Attention Decomposition

2025-09-24 · Edmund Bu, Yossi Gandelsman arxiv

We present a novel technique for interpreting the neurons in CLIP-ResNet by decomposing their contributions to the output into individual computation paths. More specifically, we analyze all pairwise combinations of neur…

Semantic Segmentation

Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias

2020-04-26 · Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 외

Common methods for interpreting neural models in natural language processing typically examine either their structure or their behavior, but not both. We propose a methodology grounded in the theory of causal mediation a…

Interpreting the neural code with Formal Concept Analysis

2008-12-01 · NeurIPS 2008 12 · Dominik Endres, Peter Foldiak

We propose a novel application of Formal Concept Analysis (FCA) to neural decoding: instead of just trying to figure out which stimulus was presented, we demonstrate how to explore the semantic relationships between the …

Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders

2025-07-21 · Krishna Kanth Nakka arxiv

Interpretability is critical in high-stakes domains such as medical imaging, where understanding model decisions is essential for clinical adoption. In this work, we introduce Sparse Autoencoder (SAE)-based interpretabil…

Rank-Order N-of-M Codes for Sparse Distributed Memory: Disentangling Representation and Learning Effects in Noise Robustness Against Contemporary Neuromorphic Architectures

2026-07-03 · Joy Bose arxiv

Large language models remain limited as continual learning systems, motivating renewed interest in Sparse Distributed Memory (SDM) as an explicit online episodic memory. CALM (Nechesov and Ruponen, 2025) identifies its t…

Continual Learning