ProtAlign: Contrastive learning paradigm for Sequence and structure alignment
Protein language models often take into consideration the alignment between a protein sequence and its textual description. However, they do not take structural information into consideration. Traditional methods treat sequence and structure separately, limiting the ability to exploit the alignment between the structure and protein sequence embeddings. In this paper, we introduce a sequence structure contrastive alignment framework, which learns a shared embedding space where proteins are represented consistently across modalities. By training on large-scale pairs of sequences and experimentally resolved or predicted structures, the model maximizes agreement between matched sequence structure pairs while pushing apart unrelated pairs. This alignment enables cross-modal retrieval (e.g., finding structural neighbors given a sequence), improves downstream prediction tasks such as function annotation and stability estimation, and provides interpretable links between sequence variation and structural organization. Our results demonstrate that contrastive learning can serve as a powerful bridge between protein sequences and structures, offering a unified representation for understanding and engineering proteins.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Modal RetrievalContrastive LearningSimilar Papers 제목 키워드 기반
Property-driven Protein Inverse Folding With Multi-Objective Preference Alignment
Protein sequence design must balance designability, defined as the ability to recover a target backbone, with multiple, often competing, developability properties such as solubility, thermostability, and expression. Exis…
Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-Sequence
In few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their applic…
Action RecognitionContrastive LearningFew-Shot action recognitionFew Shot Action Recognition+1Denoising-Contrastive Alignment for Continuous Sign Language Recognition
Continuous sign language recognition (CSLR) aims to recognize signs in untrimmed sign language videos to textual glosses. A key challenge of CSLR is achieving effective cross-modality alignment between video and gloss se…
DenoisingRepresentation LearningSign Language RecognitionContrastive learning of T cell receptor representations
Computational prediction of the interaction of T cell receptors (TCRs) and their ligands is a grand challenge in immunology. Despite advances in high-throughput assays, specificity-labelled TCR data remains sparse. In ot…
Contrastive LearningLanguage ModelingLanguage ModellingSpecificity+1Bridging Sequence-Structure Alignment in RNA Foundation Models
The alignment between RNA sequences and structures in foundation models (FMs) has yet to be thoroughly investigated. Existing FMs have struggled to establish sequence-structure alignment, hindering the free flow of genom…