paper-with-me

홈 › Papers

Explore the Limits of Omni-modal Pretraining at Scale

2024-06-13 · Yiyuan Zhang, Handong Li, Jing Liu, Xiangyu Yue

We propose to build omni-modal intelligence, which is capable of understanding any modality and learning universal representations. In specific, we propose a scalable pretraining paradigm, named Multimodal Context (MiCo), which can scale up the numbers of modalities and amount of data, together with the model parameters, in the pretraining process. With MiCo, the pretrained models show significant emergent abilities in multimodal learning, which are evaluated on the following tasks: i) single-modality perception benchmarks of 10 different modalities, ii) 25 cross-modality understanding tasks of retrieval, question-answering, captioning, and iii) 18 multimodal large language model benchmarks. Our models establish 37 new records for state-of-the-art performance. We hope that our research could contribute to the development of omni-modal intelligence. Code and Models are at https://github.com/invictus717/MiCo

📄 PDF Abstract BibTeX arXiv:2406.09412

Code (1)

invictus717/MiCo 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelQuestion AnsweringRetrievalVideo RetrievalVisual Question Answering

Similar Papers 제목 키워드 기반

OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation

2025-09-18 · Bo-Wen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen 외 arxiv

Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexp…

Representation LearningSemantic Segmentation

BrainOmni: A Brain Foundation Model for Unified EEG and MEG Signals

2025-05-18 · Qinfan Xiao, Ziyun Cui, Chi Zhang, Siqi Chen 외

Electroencephalography (EEG) and magnetoencephalography (MEG) measure neural activity non-invasively by capturing electromagnetic fields generated by dendritic currents. Although rooted in the same biophysics, EEG and ME…

EEG

OmniSat: Self-Supervised Modality Fusion for Earth Observation

2024-04-12 · Guillaume Astruc, Nicolas Gonthier, Clement Mallet, Loic Landrieu

The diversity and complementarity of sensors available for Earth Observations (EO) calls for developing bespoke self-supervised multimodal learning approaches. However, current multimodal EO datasets and models typically…

DiversityEarth ObservationLand Cover ClassificationSelf-Supervised Learning

VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

2023-04-17 · Jing Liu, Sihan Chen, Xingjian He, Longteng Guo 외

In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation. Different from widely-studied vision-language pretraining models, VALOR jointly mo…

Audio captioningAudio-Video Question Answering (AVQA)Audio-Visual CaptioningAudio-visual Question Answering+16

Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder

2026-02-27 · Kin Wai Lau, Yasar Abbas Ur Rehman, Lai-Man Po, Pedro Porto Buarque de Gusmão arxiv

Recent multimodal systems often rely on separate expert modality encoders which cause linearly scaling complexity and computational overhead with added modalities. While unified Omni-models address this via Mixture-of-Ex…