paper-with-me

Papers

Sparse Autoencoders for Low-$N$ Protein Function Prediction and Design

2025-08-25 · Darin Tsui, Kunal Talreja, Amirali Aghazadeh arxiv

Predicting protein function from amino acid sequence remains a central challenge in data-scarce (low-$N$) regimes, limiting machine learning-guided protein design when only small amounts of assay-labeled sequence-function data are available. Protein language models (pLMs) have advanced the field by providing evolutionary-informed embeddings and sparse autoencoders (SAEs) have enabled decomposition of these embeddings into interpretable latent variables that capture structural and functional features. However, the effectiveness of SAEs for low-$N$ function prediction and protein design has not been systematically studied. Herein, we evaluate SAEs trained on fine-tuned ESM2 embeddings across diverse fitness extrapolation and protein engineering tasks. We show that SAEs, with as few as 24 sequences, consistently outperform or compete with their ESM2 baselines in fitness prediction, indicating that their sparse latent space encodes compact and biologically meaningful representations that generalize more effectively from limited data. Moreover, steering predictive latents exploits biological motifs in pLM representations, yielding top-fitness variants in 83% of cases compared to designing with ESM2 alone.

📄 PDF Abstract BibTeX arXiv:2508.18567

Code (0)

등록된 구현이 없습니다.

Tasks

Protein Function PredictionProtein Design

Similar Papers 제목 키워드 기반

VFUSE: Virulent Feature Understanding with Sparse autoEncoders

2026-06-08 · Michael Yu, Matthew L. Olson arxiv

Generative models have shown remarkable progress in a variety of domains such as protein design, but such power enables the opaque generation of hazardous proteins. In this work, we introduce VFUSE (Virulent Feature Unde…

Protein Design

Circuit Tracing in Autoregressive Protein Language Models

2026-06-14 · Darin Tsui, William Deinzer, Daniel Saeedi, Amirali Aghazadeh arxiv

Protein language models (pLMs) can generate novel protein sequences with properties beyond those observed in nature, yet the mechanisms underlying protein generation remain poorly understood. Existing mechanistic interpr…

Representation Learning

Towards Interpretable Protein Structure Prediction with Sparse Autoencoders

2025-03-11 · Nithin Parsan, David J. Yang, John J. Yang

Protein language models have revolutionized structure prediction, but their nonlinear nature obscures how sequence representations inform structure prediction. While sparse autoencoders (SAEs) offer a path to interpretab…

PredictionProtein Structure Prediction

InterPLM: Discovering Interpretable Features in Protein Language Models via Sparse Autoencoders

2024-11-13 · Elana Simon, James Zou

Protein language models (PLMs) have demonstrated remarkable success in protein modeling and design, yet their internal mechanisms for predicting structure and function remain poorly understood. Here we present a systemat…

Interpreting and Steering Protein Language Models through Sparse Autoencoders

2025-02-13 · Edith Natalia Villegas Garcia, Alessio Ansuini

The rapid advancements in transformer-based language models have revolutionized natural language processing, yet understanding the internal mechanisms of these models remains a significant challenge. This paper explores …