paper-with-me

홈 › Papers

The Origins of Representation Manifolds in Large Language Models

2025-05-23 · Alexander Modell, Patrick Rubin-Delanchy, Nick Whiteley

There is a large ongoing scientific effort in mechanistic interpretability to map embeddings and internal representations of AI systems into human-understandable concepts. A key element of this effort is the linear representation hypothesis, which posits that neural representations are sparse linear combinations of `almost-orthogonal' direction vectors, reflecting the presence or absence of different features. This model underpins the use of sparse autoencoders to recover features from representations. Moving towards a fuller model of features, in which neural representations could encode not just the presence but also a potentially continuous and multidimensional value for a feature, has been a subject of intense recent discourse. We describe why and how a feature might be represented as a manifold, demonstrating in particular that cosine similarity in representation space may encode the intrinsic geometry of a feature through shortest, on-manifold paths, potentially answering the question of how distance in representation space and relatedness in concept space could be connected. The critical assumptions and predictions of the theory are validated on text embeddings and token activations of large language models.

📄 PDF Abstract BibTeX arXiv:2505.18235

Code (1)

alexandermodell/representation-manifolds 공식 구현

Similar Papers 제목 키워드 기반

On the Origins of Linear Representations in Large Language Models

2024-03-06 · Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon Aragam 외

Recent works have argued that high-level semantic concepts are encoded "linearly" in the representation space of large language models. In this work, we study the origins of such linear representations. To that end, we i…

Language ModelingLanguage ModellingLarge Language Model

Emergence of Separable Manifolds in Deep Language Representations

2020-06-01 · ICML 2020 1 · Jonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson 외

Deep neural networks (DNNs) have shown much empirical success in solving perceptual tasks across various cognitive modalities. While they are only loosely inspired by the biological brain, recent studies report considera…

Grokking From Abstraction to Intelligence

2026-03-31 · Junjie Zhang, Zhen Shen, Gang Xiong, Xisong Dong arxiv

Grokking in modular arithmetic has established itself as the quintessential fruit fly experiment, serving as a critical domain for investigating the mechanistic origins of model generalization. Despite its significance, …

REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model

2025-09-26 · Bo Li, Guanzhi Deng, Ronghao Chen, Junrong Yue 외 arxiv

Understanding how Large Language Models (LLMs) perform complex reasoning and their failure mechanisms is a challenge in interpretability research. To provide a measurable geometric analysis perspective, we define the con…

Learning Low-dimensional Manifolds for Scoring of Tissue Microarray Images

2021-02-22 · Donghui Yan, Jian Zou, Zhenpeng Li

Tissue microarray (TMA) images have emerged as an important high-throughput tool for cancer study and the validation of biomarkers. Efforts have been dedicated to further improve the accuracy of TACOMA, a cutting-edge au…

Representation Learning