UniCorn: A Unified Contrastive Learning Approach for Multi-view Molecular Representation Learning
Recently, a noticeable trend has emerged in developing pre-trained foundation models in the domains of CV and NLP. However, for molecular pre-training, there lacks a universal model capable of effectively applying to various categories of molecular tasks, since existing prevalent pre-training methods exhibit effectiveness for specific types of downstream tasks. Furthermore, the lack of profound understanding of existing pre-training methods, including 2D graph masking, 2D-3D contrastive learning, and 3D denoising, hampers the advancement of molecular foundation models. In this work, we provide a unified comprehension of existing pre-training methods through the lens of contrastive learning. Thus their distinctions lie in clustering different views of molecules, which is shown beneficial to specific downstream tasks. To achieve a complete and general-purpose molecular representation, we propose a novel pre-training framework, named UniCorn, that inherits the merits of the three methods, depicting molecular views in three different levels. SOTA performance across quantum, physicochemical, and biological tasks, along with comprehensive ablation study, validate the universality and effectiveness of UniCorn.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningDenoisingmolecular representationRepresentation LearningSimilar Papers 제목 키워드 기반
Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration
Data matching - which decides whether two data elements (e.g., string, tuple, column, or knowledge graph entity) are the "same" (a.k.a. a match) - is a key concept in data integration, such as entity matching and schema …
Data IntegrationEntity ResolutionMixture-of-ExpertsZero-Shot LearningTowards Grand Unification of Object Tracking
We present a unified method, termed Unicorn, that can simultaneously solve four tracking problems (SOT, MOT, VOS, MOTS) with a single network using the same model parameters. Due to the fragmented definitions of the obje…
Multi-Object TrackingMulti-Object Tracking and SegmentationMultiple Object TrackingObject+3Unicorn: Unified Neural Image Compression with One Number Reconstruction
Prevalent lossy image compression schemes can be divided into: 1) explicit image compression (EIC), including traditional standards and neural end-to-end algorithms; 2) implicit image compression (IIC) based on implicit …
DecoderImage CompressionUniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
While Unified Multimodal Models (UMMs) have achieved remarkable success in cross-modal comprehension, a significant gap persists in their ability to leverage such internal knowledge for high-quality generation. We formal…
Image GenerationMV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Model
Human expertise in chemistry and biomedicine relies on contextual molecular understanding, a capability that large language models (LLMs) can extend through fine-grained alignment between molecular structures and text. R…
cross-modal alignmentLanguage ModelingLanguage Modelling