paper-with-me

홈 › Papers

From superposition to sparse codes: interpretable representations in neural networks

2025-03-03 · David Klindt, Charles O'Neill, Patrik Reizinger, Harald Maurer, Nina Miolane

Understanding how information is represented in neural networks is a fundamental challenge in both neuroscience and artificial intelligence. Despite their nonlinear architectures, recent evidence suggests that neural networks encode features in superposition, meaning that input concepts are linearly overlaid within the network's representations. We present a perspective that explains this phenomenon and provides a foundation for extracting interpretable representations from neural activations. Our theoretical framework consists of three steps: (1) Identifiability theory shows that neural networks trained for classification recover latent features up to a linear transformation. (2) Sparse coding methods can extract disentangled features from these representations by leveraging principles from compressed sensing. (3) Quantitative interpretability metrics provide a means to assess the success of these methods, ensuring that extracted features align with human-interpretable concepts. By bridging insights from theoretical neuroscience, representation learning, and interpretability research, we propose an emerging perspective on understanding neural representations in both artificial and biological systems. Our arguments have implications for neural coding theories, AI transparency, and the broader goal of making deep learning models more interpretable.

📄 PDF Abstract BibTeX arXiv:2503.01824

Code (0)

등록된 구현이 없습니다.

Tasks

compressed sensingRepresentation Learning

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Capacity-achieving Sparse Superposition Codes via Approximate Message Passing Decoding

2015-01-23 · Cynthia Rush, Adam Greig, Ramji Venkataramanan

Sparse superposition codes were recently introduced by Barron and Joseph for reliable communication over the AWGN channel at rates approaching the channel capacity. The codebook is defined in terms of a Gaussian design m…

Decoder

Superposition disentanglement of neural representations reveals hidden alignment

2025-10-03 · André Longon, David Klindt, Meenakshi Khosla arxiv

The superposition hypothesis states that single neurons may participate in representing multiple features in order for the neural network to represent more features than it has neurons. In neuroscience and AI, representa…

Compressed Computation under $L^4$ Loss is likely Computation in Superposition

2026-07-06 · Francisco Ferreira da Silva, Stefan Heimersheim arxiv

Neural networks are thought to represent concepts as directions in their activation space, and superposition lets them encode more concepts than they have dimensions. It is natural to ask whether they can also compute mo…

Can sparse autoencoders make sense of latent representations?

2024-10-15 · Viktoria Schuster

Sparse autoencoders (SAEs) have lately been used to uncover interpretable latent features in large language models. Here, we explore their potential for decomposing latent representations in complex and high-dimensional …

Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning

2024-10-31 · John Wu, David Wu, Jimeng Sun

Medical coding, the translation of unstructured clinical text into standardized medical codes, is a crucial but time-consuming healthcare practice. Though large language models (LLM) could automate the coding process and…

Dictionary LearningLanguage ModelingLanguage Modelling