paper-with-me

홈 › Papers

Sparse autoencoders reveal organized biological knowledge but minimal regulatory logic in single-cell foundation models: a comparative atlas of Geneformer and scGPT

2026-03-03 · Ihor Kendiukhov arxiv

Background: Single-cell foundation models such as Geneformer and scGPT encode rich biological information, but whether this includes causal regulatory logic rather than statistical co-expression remains unclear. Sparse autoencoders (SAEs) can resolve superposition in neural networks by decomposing dense activations into interpretable features, yet they have not been systematically applied to biological foundation models. Results: We trained TopK SAEs on residual stream activations from all layers of Geneformer V2-316M (18 layers, d=1152) and scGPT whole-human (12 layers, d=512), producing atlases of 82525 and 24527 features, respectively. Both atlases confirm massive superposition, with 99.8 percent of features invisible to SVD. Systematic characterization reveals rich biological organization: 29 to 59 percent of features annotate to Gene Ontology, KEGG, Reactome, STRING, or TRRUST, with U-shaped layer profiles reflecting hierarchical abstraction. Features organize into co-activation modules (141 in Geneformer, 76 in scGPT), exhibit causal specificity (median 2.36x), and form cross-layer information highways (63 to 99.8 percent). When tested against genome-scale CRISPRi perturbation data, only 3 of 48 transcription factors (6.2 percent) show regulatory-target-specific feature responses. A multi-tissue control yields marginal improvement (10.4 percent, 5 of 48 TFs), establishing model representations as the bottleneck. Conclusions: These models have internalized organized biological knowledge, including pathway membership, protein interactions, functional modules, and hierarchical abstraction, yet they encode minimal causal regulatory logic. We release both feature atlases as interactive web platforms enabling exploration of more than 107000 features across 30 layers of two leading single-cell foundation models.

📄 PDF Abstract BibTeX arXiv:2603.02952

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Domain-Filtered Knowledge Graphs from Sparse Autoencoder Features

2026-04-26 · John Winnicki, Abeynaya Gnanasekaran, Eric Darve arxiv

Sparse autoencoders (SAEs) extract millions of interpretable features from a language model, but flat feature inventories aren't very useful on their own. Domain concepts get mixed with generic and weakly grounded featur…

Knowledge Graphs

A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders

2026-06-05 · Chenhao Zhang, Chris Lin, Su-In Lee arxiv

We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs). While SAEs improve interpretability of neural networks by learning sp…

Towards Interpretable Protein Structure Prediction with Sparse Autoencoders

2025-03-11 · Nithin Parsan, David J. Yang, John J. Yang

Protein language models have revolutionized structure prediction, but their nonlinear nature obscures how sequence representations inform structure prediction. While sparse autoencoders (SAEs) offer a path to interpretab…

PredictionProtein Structure Prediction

Rectified Factor Networks

2015-02-23 · NeurIPS 2015 12 · Djork-Arné Clevert, Andreas Mayr, Thomas Unterthiner, Sepp Hochreiter

We propose rectified factor networks (RFNs) to efficiently construct very sparse, non-linear, high-dimensional representations of the input. RFN models identify rare and small events in the input, have a low interference…

Drug Discovery

Learning biologically relevant features in a pathology foundation model using sparse autoencoders

2024-07-15 · Nhat Minh Le, Ciyue Shen, Neel Patel, Chintan Shah 외

Pathology plays an important role in disease diagnosis, treatment decision-making and drug development. Previous works on interpretability for machine learning models on pathology images have revolved around methods such…