paper-with-me

Papers

Adaptive Protein Tokenization

2026-02-06 · Rohit Dilip, Ayush Varshney, David Van Valen arxiv

Tokenization is a promising path to multi-modal models capable of jointly understanding protein sequences, structure, and function. Existing protein structure tokenizers create tokens by pooling information from local neighborhoods, an approach that limits their performance on generative and representation tasks. In this work, we present a method for global tokenization of protein structures in which successive tokens contribute increasing levels of detail to a global representation. This change resolves several issues with generative models based on local protein tokenization: it mitigates error accumulation, provides embeddings without sequence-reduction operations, and allows task-specific adaptation of a tokenized sequence's information content. We validate our method on reconstruction, generative, and representation tasks and demonstrate that it matches or outperforms existing models based on local protein structure tokenizers. We show how adaptive tokens enable inference criteria based on information content, which boosts designability. We validate representations generated from our tokenizer on CATH classification tasks and demonstrate that non-linear probing on our tokenized sequences outperforms equivalent probing on representations from other tokenizers. Finally, we demonstrate how our method supports zero-shot protein shrinking and affinity maturation.

📄 PDF Abstract BibTeX arXiv:2602.06418

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Linguistic Laws Meet Protein Sequences: A Comparative Analysis of Subword Tokenization Methods

2024-11-26 · Burak Suyunu, Enes Taylan, Arzucan Özgür

Tokenization is a crucial step in processing protein sequences for machine learning models, as proteins are complex sequences of amino acids that require meaningful segmentation to capture their functional and structural…

evoBPE: Evolutionary Protein Sequence Tokenization

2025-03-11 · Burak Suyunu, Özdeniz Dolu, Arzucan Özgür

Recent advancements in computational biology have drawn compelling parallels between protein sequences and linguistic structures, highlighting the need for sophisticated tokenization methods that capture the intricate ev…

Protein Function Prediction

SaDiT: Efficient Protein Backbone Design via Latent Structural Tokenization and Diffusion Transformers

2026-02-06 · Shentong Mo, Lanqing Li arxiv

Generative models for de novo protein backbone design have achieved remarkable success in creating novel protein structures. However, these diffusion-based approaches remain computationally intensive and slower than desi…

From Static Structures to Ensembles: Studying and Harnessing Protein Structure Tokenization

2025-11-13 · Zijing Liu, Bin Feng, He Cao, Yu Li arxiv

Protein structure tokenization converts 3D structures into discrete or vectorized representations, enabling the integration of structural and sequence data. Despite many recent works on structure tokenization, the proper…

Protein Structure Tokenization: Benchmarking and New Recipe

2025-02-28 · Xinyu Yuan, Zichen Wang, Marcus Collins, Huzefa Rangwala

Recent years have witnessed a surge in the development of protein structural tokenization methods, which chunk protein 3D structures into discrete or continuous representations. Structure tokenization enables the direct …

BenchmarkingLanguage ModelingLanguage Modelling