paper-with-me

홈 › Papers

On the geometry and topology of representations: the manifolds of modular addition

2025-12-31 · Gabriela Moisescu-Pareja, Gavin McCracken, Harley Wiltzer, Vincent Létourneau, Colin Daniels, Doina Precup, Jonathan Love arxiv

The Clock and Pizza interpretations, associated with architectures differing in either uniform or learnable attention, were introduced to argue that different architectural designs can yield distinct circuits for modular addition. In this work, we show that this is not the case, and that both uniform attention and trainable attention architectures implement the same algorithm via topologically and geometrically equivalent representations. Our methodology goes beyond the interpretation of individual neurons and weights. Instead, we identify all of the neurons corresponding to each learned representation and then study the collective group of neurons as one entity. This method reveals that each learned representation is a manifold that we can study utilizing tools from topology. Based on this insight, we can statistically analyze the learned representations across hundreds of circuits to demonstrate the similarity between learned modular addition circuits that arise naturally from common deep learning paradigms.

📄 PDF Abstract BibTeX arXiv:2512.25060

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning geometry and topology via multi-chart flows

2025-05-30 · Hanlin Yu, Søren Hauberg, Marcelo Hartmann, Arto Klami 외

Real world data often lie on low-dimensional Riemannian manifolds embedded in high-dimensional spaces. This motivates learning degenerate normalizing flows that map between the ambient space and a low-dimensional latent …

Directed Graph Embeddings in Pseudo-Riemannian Manifolds

2021-06-16 · Aaron Sim, Maciej Wiatrak, Angus Brayne, Páidí Creed 외

The inductive biases of graph representation learning algorithms are often encoded in the background geometry of their embedding space. In this paper, we show that general directed graphs can be effectively represented b…

Graph Representation LearningLink PredictionRepresentation Learning

Continuous normalizing flows on manifolds

2021-03-14 · Luca Falorsi

Normalizing flows are a powerful technique for obtaining reparameterizable samples from complex multimodal distributions. Unfortunately, current approaches are only available for the most basic geometries and fall short …

Emergence of Separable Manifolds in Deep Language Representations

2020-06-01 · ICML 2020 1 · Jonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson 외

Deep neural networks (DNNs) have shown much empirical success in solving perceptual tasks across various cognitive modalities. While they are only loosely inspired by the biological brain, recent studies report considera…

Neural Ordinary Differential Equations on Manifolds

2020-06-11 · Luca Falorsi, Patrick Forré

Normalizing flows are a powerful technique for obtaining reparameterizable samples from complex multimodal distributions. Unfortunately current approaches fall short when the underlying space has a non trivial topology, …