paper-with-me

홈 › Papers

Measuring Feature Sparsity in Language Models

2023-10-11 · Mingyang Deng, Lucas Tao, Joe Benton

Recent works have proposed that activations in language models can be modelled as sparse linear combinations of vectors corresponding to features of input text. Under this assumption, these works aimed to reconstruct feature directions using sparse coding. We develop metrics to assess the success of these sparse coding techniques and test the validity of the linearity and sparsity assumptions. We show our metrics can predict the level of sparsity on synthetic sparse linear activations, and can distinguish between sparse linear data and several other distributions. We use our metrics to measure levels of sparsity in several language models. We find evidence that language model activations can be accurately modelled by sparse linear combinations of features, significantly more so than control datasets. We also show that model activations appear to be sparsest in the first and final layers.

📄 PDF Abstract BibTeX arXiv:2310.07837

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Sparsity and Superposition in Mixture of Experts

2025-10-26 · Marmik Chaudhari, Jeremi Nuer, Rome Thorstenson arxiv

Mixture of Experts (MoE) models have become central to scaling large language models, yet their mechanistic differences from dense networks remain poorly understood. Previous work has explored how dense models use \texti…

Feature Importance

Sparsity-Probe: Analysis tool for Deep Learning Models

2021-05-14 · Ido Ben-Shaul, Shai Dekel

We propose a probe for the analysis of deep learning architectures that is based on machine learning and approximation theoretical principles. Given a deep learning architecture and a training set, during or after traini…

BIG-bench Machine LearningDeep Learning

2SSP: A Two-Stage Framework for Structured Pruning of LLMs

2025-01-29 · Fabrizio Sandri, Elia Cunegatti, Giovanni Iacca

We propose a novel Two-Stage framework for Structured Pruning (2SSP) for pruning Large Language Models (LLMs), which combines two different strategies of pruning, namely Width and Depth Pruning. The first stage (Width Pr…

Language ModelingLanguage Modelling

SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs

2025-11-10 · Sean P. Fillingham, Andrew Gordon, Peter Lai, Xavier Poncini 외 arxiv

Mechanistic interpretability aims to decompose neural networks into interpretable features and map their connecting circuits. The standard approach trains sparse autoencoders (SAEs) on each layer's activations. However, …

Investigating Sensitive Directions in GPT-2: An Improved Baseline and Comparative Analysis of SAEs

2024-10-16 · Daniel J. Lee, Stefan Heimersheim

Sensitive directions experiments attempt to understand the computational features of Language Models (LMs) by measuring how much the next token prediction probabilities change by perturbing activations along specific dir…