Protein Language Model Embeddings Improve Generalization of Implicit Transfer Operators
Molecular dynamics (MD) is a central computational tool in physics, chemistry, and biology, enabling quantitative prediction of experimental observables as expectations over high-dimensional molecular distributions such as Boltzmann distributions and transition densities. However, conventional MD is fundamentally limited by the high computational cost required to generate independent samples. Generative molecular dynamics (GenMD) has recently emerged as an alternative, learning surrogates of molecular distributions either from data or through interaction with energy models. While these methods enable efficient sampling, their transferability across molecular systems is often limited. In this work, we show that incorporating auxiliary sources of information can improve the data efficiency and generalization of transferable implicit transfer operators (TITO) for molecular dynamics. We find that coarse-grained TITO models are substantially more data-efficient than Boltzmann Emulators, and that incorporating protein language model (pLM) embeddings further improves out-of-distribution generalization. Our approach, PLaTITO, achieves state-of-the-art performance on equilibrium sampling benchmarks for out-of-distribution protein systems, including fast-folding proteins. We further study the impact of additional conditioning signals such as structural embeddings, temperature, and large-language-model-derived embeddings on model performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Protein Language ModelSimilar Papers 제목 키워드 기반
Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction
Recently, there has been a growing adoption of protein language models (PLMs) in biomedical science. Their embeddings provide a rich numerical representation of protein sequences which achieve state-of-the-art performanc…
Protein Language ModelBioBridge: Bridging Proteins and Language for Enhanced Biological Reasoning with LLMs
Existing Protein Language Models (PLMs) often suffer from limited adaptability to multiple tasks and exhibit poor generalization across diverse biological contexts. In contrast, general-purpose Large Language Models (LLM…
Continual PretrainingProtein Language Model-Powered 3D Ligand Binding Site Prediction from Protein Sequence
Prediction of ligand binding sites of proteins is a fundamental and important task for understanding the function of proteins and screening potential drugs. Most existing methods require experimentally determined protein…
Drug DiscoveryGraph Neural NetworkLanguage ModelingLanguage Modelling+1FoldExplorer: Fast and Accurate Protein Structure Search with Sequence-Enhanced Graph Embedding
The advent of highly accurate protein structure prediction methods has fueled an exponential expansion of the protein structure database. Consequently, there is a rising demand for rapid and precise structural homolog se…
Graph AttentionGraph EmbeddingProtein Structure PredictionIntegrating protein sequence embeddings with structure via graph-based deep learning for the prediction of single-residue properties
Understanding the intertwined contributions of amino acid sequence and spatial structure is essential to explain protein behaviour. Here, we introduce INFUSSE (Integrated Network Framework Unifying Structure and Sequence…
Large Language Model