paper-with-me

홈 › Papers

Interpreting Neural Networks through the Polytope Lens

2022-11-22 · Sid Black, Lee Sharkey, Leo Grinsztajn, Eric Winsor, Dan Braun, Jacob Merizian, Kip Parker, Carlos Ramón Guevara, Beren Millidge, Gabriel Alfour, Connor Leahy

Mechanistic interpretability aims to explain what a neural network has learned at a nuts-and-bolts level. What are the fundamental primitives of neural network representations? Previous mechanistic descriptions have used individual neurons or their linear combinations to understand the representations a network has learned. But there are clues that neurons and their linear combinations are not the correct fundamental units of description: directions cannot describe how neural networks use nonlinearities to structure their representations. Moreover, many instances of individual neurons and their combinations are polysemantic (i.e. they have multiple unrelated meanings). Polysemanticity makes interpreting the network in terms of neurons or directions challenging since we can no longer assign a specific feature to a neural unit. In order to find a basic unit of description that does not suffer from these problems, we zoom in beyond just directions to study the way that piecewise linear activation functions (such as ReLU) partition the activation space into numerous discrete polytopes. We call this perspective the polytope lens. The polytope lens makes concrete predictions about the behavior of neural networks, which we evaluate through experiments on both convolutional image classifiers and language models. Specifically, we show that polytopes can be used to identify monosemantic regions of activation space (while directions are not in general monosemantic) and that the density of polytope boundaries reflect semantic boundaries. We also outline a vision for what mechanistic interpretability might look like through the polytope lens.

📄 PDF Abstract BibTeX arXiv:2211.12312

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

2026-05-30 · Hwiyeong Lee, Ingyu Bang, Uiji Hwang, Hyelim Lim 외 arxiv

While sparse autoencoders provide features more interpretable than individual neurons, reliably characterizing them remains challenging. We propose Query Lens, which extends Logit Lens to enable more comprehensive and fa…

AffineLens: Capturing the Continuous Piecewise Affine Functions of Neural Networks

2026-05-07 · Yi Wei, Xuan Qi, Furao Shen, Jian Zhao 외 arxiv

Piecewise affine neural networks (PANNs) provide a principled geometric perspective on neural network expressivity by characterizing the input--output map as a continuous piecewise affine (CPA) function whose complexity …

Linear Convergence of the Frank-Wolfe Algorithm over Product Polytopes

2025-05-16 · Gabriele Iommazzo, David Martínez-Rubio, Francisco Criado, Elias Wirth 외

We study the linear convergence of Frank-Wolfe algorithms over product polytopes. We analyze two condition numbers for the product polytope, namely the \emph{pyramidal width} and the \emph{vertex-facet distance}, based o…

Attention Lens: A Tool for Mechanistically Interpreting the Attention Head Information Retrieval Mechanism

2023-10-25 · Mansi Sakarvadia, Arham Khan, Aswathy Ajith, Daniel Grzenda 외

Transformer-based Large Language Models (LLMs) are the state-of-the-art for natural language tasks. Recent work has attempted to decode, by reverse engineering the role of linear layers, the internal mechanisms by which …

Information RetrievalRetrieval

On Identifiable Polytope Characterization for Polytopic Matrix Factorization

2022-04-25 · Bariscan Bozkurt, Alper T. Erdogan

Polytopic matrix factorization (PMF) is a recently introduced matrix decomposition method in which the data vectors are modeled as linear transformations of samples from a polytope. The successful recovery of the origina…