paper-with-me

홈 › Papers

Toy Models of Superposition

2022-09-21 · Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, Christopher Olah

Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This paper provides a toy model where polysemanticity can be fully understood, arising as a result of models storing additional sparse features in "superposition." We demonstrate the existence of a phase change, a surprising connection to the geometry of uniform polytopes, and evidence of a link to adversarial examples. We also discuss potential implications for mechanistic interpretability.

📄 PDF Abstract BibTeX arXiv:2209.10652

Code (1)

anthropics/toy-models-of-superposition 공식 구현

Similar Papers 제목 키워드 기반

SCL(FOL) Can Simulate Non-Redundant Superposition Clause Learning

2023-05-22 · Martin Bromberger, Chaahat Jain, Christoph Weidenbach

We show that SCL(FOL) can simulate the derivation of non-redundant clauses by superposition for first-order logic without equality. Superposition-based reasoning is performed with respect to a fixed reduction ordering. T…

Superposition unifies power-law training dynamics

2026-02-01 · Zixin Jessie Chen, Hao Chen, Yizhou Liu, Jeff Gore arxiv

We investigate the role of feature superposition in the emergence of power-law training dynamics using a teacher-student framework. We first derive an analytic theory for training without superposition, establishing that…

Mathematical Models of Computation in Superposition

2024-08-10 · Kaarel Hänni, Jake Mendel, Dmitry Vaintrob, Lawrence Chan

Superposition -- when a neural network represents more ``features'' than it has dimensions -- seems to pose a serious challenge to mechanistically interpreting current AI systems. Existing theory work studies \emph{repre…

Adversarial Examples Are Not Bugs, They Are Superposition

2025-08-24 · Liv Gorton, Owen Lewis arxiv

Adversarial examples -- inputs with imperceptible perturbations that fool neural networks -- remain one of deep learning's most perplexing phenomena despite nearly a decade of research. While numerous defenses and explan…

Few-shot Named Entity Recognition via Superposition Concept Discrimination

2024-03-25 · Jiawei Chen, Hongyu Lin, Xianpei Han, Yaojie Lu 외

Few-shot NER aims to identify entities of target types with only limited number of illustrative instances. Unfortunately, few-shot NER is severely challenged by the intrinsic precise generalization problem, i.e., it is h…

Active Learningfew-shot-nerFew-shot NERnamed-entity-recognition+2