paper-with-me

Papers

Training-Free Generation of Protein Sequences from Small Family Alignments via Stochastic Attention

2026-03-16 · Jeffrey D. Varner arxiv

Generating novel protein sequences that respect a family's statistical constraints typically requires training deep generative models on thousands to millions of examples. Yet most protein families are small: the median Pfam seed alignment contains only 22 sequences, a regime where learned models overfit or collapse. We propose \emph{stochastic attention} (SA), a training-free sampler that treats the modern Hopfield energy over stored sequences as a Boltzmann distribution and draws samples via Langevin dynamics. The score function is the residual of a single softmax attention operation, eliminating the need for a trained score network, pretraining data, or graphics processing units (GPUs). Across eight Pfam families spanning 37 to 420 sequences and 23 to 262 residues, SA generates sequences with low composition divergence, novelty, and structural plausibility supported by ESMFold and AlphaFold2. Compared with profile hidden Markov models (HMMs), EvoDiff, and the multiple sequence alignment (MSA) Transformer, SA is the only tested method to simultaneously achieve low composition divergence, genuine novelty, and sequence identity within each family's nearest-neighbor identity range; the others drift outside this range or produce near-copies. The critical inverse temperature is predicted from principal component analysis (PCA) dimensionality alone, enabling fully automatic operation from a seed alignment. In two domains with deep mutational scanning data, SA-generated substitutions are enriched for experimentally tolerated mutations beyond a position-matched null, and an independent language model (ESM2-650M) scores them within the natural range. Stochastic attention thus opens training-free sequence generation to the long tail of protein families too small for deep learning.

📄 PDF Abstract BibTeX arXiv:2603.14717

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple Sequence Alignment

Similar Papers 제목 키워드 기반

Generative power of a protein language model trained on multiple sequence alignments

2022-04-14 · Damiano Sgarbossa, Umberto Lupo, Anne-Florence Bitbol

Computational models starting from large ensembles of evolutionarily related protein sequences capture a representation of protein families and learn constraints associated to protein structure and function. They thus op…

Language ModelingLanguage ModellingMasked Language ModelingProtein Design+1

All-atom inverse protein folding through discrete flow matching

2025-07-04 · Kai Yi, Kiarash Jamali, Sjors H. W. Scheres arxiv

The recent breakthrough of AlphaFold3 in modeling complex biomolecular interactions, including those between proteins and ligands, nucleotides, or metal ions, creates new opportunities for protein design. In so-called in…

Protein Design

Few Shot Protein Generation

2022-04-03 · Soumya Ram, Tristan Bepler

We present the MSA-to-protein transformer, a generative model of protein sequences conditioned on protein families represented by multiple sequence alignments (MSAs). Unlike existing approaches to learning generative mod…

Multiple Sequence Alignment

DiffDTM: A conditional structure-free framework for bioactive molecules generation targeted for dual proteins

2023-06-24 · Lei Huang, Zheng Yuan, Huihui Yan, Rong Sheng 외

Advances in deep generative models shed light on de novo molecule generation with desired properties. However, molecule generation targeted for dual protein targets still faces formidable challenges including protein 3D …

Design Proteins Using Large Language Models: Enhancements and Comparative Analyses

2024-08-12 · Kamyar Zeinalipour, Neda Jamshidi, Monica Bianchini, Marco Maggini 외

Pre-trained LLMs have demonstrated substantial capabilities across a range of conventional natural language processing (NLP) tasks, such as summarization and entity recognition. In this paper, we explore the application …

valid