paper-with-me

홈 › Papers

Linear Explanations for Individual Neurons

2024-05-10 · Tuomas Oikarinen, Tsui-Wei Weng

In recent years many methods have been developed to understand the internal workings of neural networks, often by describing the function of individual neurons in the model. However, these methods typically only focus on explaining the very highest activations of a neuron. In this paper we show this is not sufficient, and that the highest activation range is only responsible for a very small percentage of the neuron's causal effect. In addition, inputs causing lower activations are often very different and can't be reliably predicted by only looking at high activations. We propose that neurons should instead be understood as a linear combination of concepts, and develop an efficient method for producing these linear explanations. In addition, we show how to automatically evaluate description quality using simulation, i.e. predicting neuron activations on unseen inputs in vision setting.

📄 PDF Abstract BibTeX arXiv:2405.06855

Code (1)

Trustworthy-ML-Lab/Linear-Explanations 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Rigorously Assessing Natural Language Explanations of Neurons

2023-09-19 · Jing Huang, Atticus Geiger, Karel D'Oosterlinck, Zhengxuan Wu 외

Natural language is an appealing medium for explaining how large language models process and store information, but evaluating the faithfulness of such explanations is challenging. To help address this, we develop two mo…

Influence-Directed Explanations for Deep Convolutional Networks

2018-02-11 · ICLR 2018 1 · Klas Leino, Shayak Sen, Anupam Datta, Matt Fredrikson 외

We study the problem of explaining a rich class of behavioral properties of deep neural networks. Distinctively, our influence-directed explanations approach this problem by peering inside the network to identify neurons…

Global Concept-Based Interpretability for Graph Neural Networks via Neuron Analysis

2022-08-22 · Han Xuanyuan, Pietro Barbiero, Dobrik Georgiev, Lucie Charlotte Magister 외

Graph neural networks (GNNs) are highly effective on a variety of graph-related tasks; however, they lack interpretability and transparency. Current explainability approaches are typically local and treat GNNs as black-b…

NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models

2026-01-07 · Weiqi Liu, Yongliang Miao, Haiyan Zhao, Yanguang Liu 외 arxiv

Neuron-level interpretation in large language models (LLMs) is fundamentally challenged by widespread polysemanticity, where individual neurons respond to multiple distinct semantic concepts. Existing single-pass interpr…

Deciphering Functions of Neurons in Vision-Language Models

2025-02-10 · Jiaqi Xu, Cuiling Lan, Xuejin Chen, Yan Lu

The burgeoning growth of open-sourced vision-language models (VLMs) has catalyzed a plethora of applications across diverse domains. Ensuring the transparency and interpretability of these models is critical for fosterin…