paper-with-me

홈 › Papers

Probing for Representation Manifolds in Superposition

2026-05-18 · Alexander Modell arxiv

This paper introduces the Manifold Probe, a supervised method for discovering representation manifolds in superposition. The method generalizes linear regression probes by learning the space of features of a concept that can be linearly predicted from the representations, and then learning the directions used to encode them. We demonstrate the probe on representations of time and space in Llama 2-7b, finding manifolds which linearly represent an interpretable set of features in each case. In the case of time, we show that by steering along the manifold, we can influence the model's completions about the years in which famous songs, movies and books were released, providing evidence that the Manifold Probe can discover manifolds which are causally involved in model behaviour.

📄 PDF Abstract BibTeX arXiv:2605.18537

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models

2026-04-07 · Michael Rizvi-Martel, Guillaume Rabusseau, Marius Mosbach arxiv

Latent reasoning via continuous chain-of-thoughts (Latent CoT) has emerged as a promising alternative to discrete CoT reasoning. Operating in continuous space increases expressivity and has been hypothesized to enable su…

The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors

2026-02-02 · Raphaël Sarfati, Eric Bigelow, Daniel Wurgaft, Siddharth Boppana 외 arxiv

Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these beliefs are encoded in representation space, how they update with new evi…

Geometric Limits of Knowledge Distillation: A Minimum-Width Theorem via Superposition Theory

2026-04-05 · Nilesh Sarkar, Dawar Jyoti Deka arxiv

Knowledge distillation compresses large teachers into smaller students, but performance saturates at a loss floor that persists across training methods and objectives. We argue this floor is geometric: neural networks re…

Knowledge Distillation

Probing Human Visual Robustness with Neurally-Guided Deep Neural Networks

2024-05-04 · Zhenan Shao, Linjian Ma, Yiqing Zhou, Yibo Jacky Zhang 외

Humans effortlessly navigate the dynamic visual world, yet deep neural networks (DNNs), despite excelling at many visual tasks, are surprisingly vulnerable to minor image perturbations. Past theories suggest that human v…

Decision MakingNavigateObject Recognition

Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalisation

2026-03-30 · Vitória Barin Pacela, Shruti Joshi, Isabela Camacho, Simon Lacoste-Julien 외 arxiv

The linear representation hypothesis states that neural network activations encode high-level concepts as linear mixtures. However, under superposition, this encoding is a projection from a higher-dimensional concept spa…