paper-with-me

홈 › Papers

BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language

2026-06-20 · Qizhi Pei, Zhimeng Zhou, Yi Duan, Yiyang Zhao, Wei Li, Han Guo, Liang He, Chengping Li, Chang-Yu Hsieh, Conghui He, Rui Yan, Lijun Wu arxiv

We present BioMatrix, the first multimodal foundation model that natively integrates sequences, structures, and natural language for both molecules and proteins within a single decoder-only architecture. Existing biological foundation models pursue native multimodality and broad entity coverage separately: those that fuse multiple modalities under a shared objective remain confined to a single entity type, while those spanning multiple entity types either omit explicit structural modeling or rely on adapter-based designs in which the model cannot natively generate the very modalities it can read. BioMatrix closes this gap by mapping molecular sequences (supporting both SMILES and SELFIES notations), molecular structures, protein sequences, protein structures, and natural language into a shared discrete token space through a unified tokenization scheme, so that all modalities are consumed and produced uniformly under a single next-token prediction objective -- without external encoders, projection adapters, or modality-specific output heads. Built upon the Qwen3 language model (1.7B and 4B), BioMatrix is continually pretrained on 304.4 billion tokens spanning general and domain-specific text, sequence and structure views of molecules and proteins, and cross-modal corpora that interleave biomolecular entities with scientific text and link distinct entities through molecule-protein and protein-protein interaction data. After tuning on a comprehensive suite of downstream applications covering 80 tasks across 6 categories -- encompassing single-entity and multi-entity understanding and generation tasks across and within modalities -- BioMatrix achieves state-of-the-art or competitive performance on 77 out of 80 tasks, demonstrating that a single, natively multimodal generalist model can effectively match or surpass specialized approaches across a wide range of biological tasks.

📄 PDF Abstract BibTeX arXiv:2606.22138

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning

2025-10-01 · Ching-Huei Tsou, Michal Ozery-Flato, Ella Barkan, Diwakar Mahajan 외 arxiv

Recent advances in large language models (LLMs) and biomedical foundation models (BioFMs) have achieved strong results in biological text reasoning, molecular modeling, and single-cell analysis, yet they remain siloed in…

Question Answering

Inertia-1: An Open Exploration of Wearable Motion Foundation Models

2026-07-07 · Zongzhe Xu, Aakarsh Anand, Sarah Jiang, Chuntung Zhuang 외 arxiv

Wearable motion sensing provides a continuous and scalable window into human behavior and health, making it a natural fit for foundation models, yet its pretraining and scaling principles remain poorly understood. Prior …

Human Activity RecognitionRepresentation Learning

Automatic Image-Level Morphological Trait Annotation for Organismal Images

2026-04-02 · Vardaan Pahuja, Samuel Stevens, Alyson East, Sydne Record 외 arxiv

Morphological traits are physical characteristics of biological organisms that provide vital clues on how organisms interact with their environment. Yet extracting these traits remains a slow, expert-driven process, limi…

FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning

2025-12-14 · Yue Jiang, Dingkang Yang, Minghao Han, Jinghang Han 외 arxiv

Despite rapid progress in multimodal large language models (MLLMs) and emerging omni-modal architectures, current benchmarks remain limited in scope and integration, suffering from incomplete modality coverage, restricte…

EVA: Towards a universal model of the immune system

2026-02-10 · Scienta Team, Ethan Bandasack, Vincent Bouget, Apolline Bruley 외 arxiv

The effective application of foundation models to translational research in immune-mediated diseases requires multimodal patient-level representations that can capture complex phenotypes emerging from multicellular inter…

Activity PredictionTransfer Learning