paper-with-me

홈 › Papers

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

2024-07-16 · Zehan Wang, Ziang Zhang, Hang Zhang, Luping Liu, Rongjie Huang, Xize Cheng, Hengshuang Zhao, Zhou Zhao

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representation in understanding and generation pipelines, high-quality omni joint representations would be a step toward co-processing more diverse multimodal information. In this work, we present OmniBind, large-scale multimodal joint representation models ranging in scale from 7 billion to 30 billion parameters, which support 3D, audio, image, and language inputs. Due to the scarcity of data pairs across all modalities, instead of training large models from scratch, we propose remapping and binding the spaces of various pre-trained specialist models together. This approach enables "scaling up" by indirectly increasing the model parameters and the amount of seen data. To effectively integrate various spaces, we dynamically assign weights to different spaces by learning routers with two objectives: cross-modal overall alignment and language representation decoupling. Notably, since binding and routing spaces both only require lightweight networks, OmniBind is extremely training-efficient. Learning the largest 30B model requires merely unpaired unimodal data and approximately 3 days on a single 8-4090 node. Extensive experiments demonstrate the versatility and superiority of OmniBind as an omni representation model, highlighting its great potential for diverse applications, such as any-query and composable multimodal understanding.

📄 PDF Abstract BibTeX arXiv:2407.11895

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniBind: Teach to Build Unequal-Scale Modality Interaction for Omni-Bind of All

2024-05-25 · Yuanhuiyi Lyu, Xu Zheng, Dahun Kim, Lin Wang

Research on multi-modal learning dominantly aligns the modalities in a unified space at training, and only a single one is taken for prediction at inference. However, for a real machine, e.g., a robot, sensors could be a…

Allcross-modal alignment

Kling-Omni Technical Report

2025-12-18 · Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du 외 arxiv

We present Kling-Omni, a generalist generative framework designed to synthesize high-fidelity videos directly from multimodal visual language inputs. Adopting an end-to-end perspective, Kling-Omni bridges the functional …

Instruction FollowingVideo Generation

Explore the Limits of Omni-modal Pretraining at Scale

2024-06-13 · Yiyuan Zhang, Handong Li, Jing Liu, Xiangyu Yue

We propose to build omni-modal intelligence, which is capable of understanding any modality and learning universal representations. In specific, we propose a scalable pretraining paradigm, named Multimodal Context (MiCo)…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+4

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

2025-06-16 · Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou 외

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatena…

Large Language Modelmultimodal interaction

NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching

2025-10-15 · Run Luo, Xiaobo Xia, Lu Wang, Longze Chen 외 arxiv

Next-generation multimodal foundation models capable of any-to-any cross-modal generation and multi-turn interaction will serve as core components of artificial general intelligence systems, playing a pivotal role in hum…

Cross-Modal Retrievalmultimodal generation