paper-with-me

홈 › Papers

e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings

2026-01-07 · Haonan Chen, Sicheng Gao, Radu Timofte, Tetsuya Sakai, Zhicheng Dou arxiv

Modern information systems often involve different types of items, e.g., a text query, an image, a video clip, or an audio segment. This motivates omni-modal embedding models that map heterogeneous modalities into a shared space for direct comparison. However, most recent omni-modal embeddings still rely heavily on implicit alignment inherited from pretrained vision-language model (VLM) backbones. In practice, this causes three common issues: (i) similarity logits have modality-dependent sharpness, so scores are not on a consistent scale; (ii) in-batch negatives become less effective over time because mixed-modality batches create an imbalanced hardness distribution; as a result, many negatives quickly become trivial and contribute little gradient; and (iii) embeddings across modalities show mismatched first- and second-order statistics, which makes rankings less stable. To tackle these problems, we propose e5-omni, a lightweight explicit alignment recipe that adapts off-the-shelf VLMs into robust omni-modal embedding models. e5-omni combines three simple components: (1) modality-aware temperature calibration to align similarity scales, (2) a controllable negative curriculum with debiasing to focus on confusing negatives while reducing the impact of false negatives, and (3) batch whitening with covariance regularization to better match cross-modal geometry in the shared embedding space. Experiments on MMEB-V2 and AudioCaps show consistent gains over strong bi-modal and omni-modal baselines, and the same recipe also transfers well to other VLM backbones. We release our model checkpoint at https://huggingface.co/Haon-Chen/e5-omni-7B.

📄 PDF Abstract BibTeX arXiv:2601.03666

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

2025-10-17 · Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang 외 arxiv

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM.…

OpenOmni: Large Language Models Pivot Zero-shot Omnimodal Alignment across Language with Real-time Self-Aware Emotional Speech Synthesis

2025-01-08 · Run Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu 외

Recent advancements in omnimodal learning have been achieved in understanding and generation across images, text, and speech, though mainly within proprietary models. Limited omnimodal datasets and the inherent challenge…

DecoderEmotional Speech SynthesisLanguage ModelingLanguage Modelling+2

Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs

2026-04-07 · Hongcheng Liu, Yuhao Wang, Zhe Chen, Pingjie Wang 외 arxiv

Omni Large Language Models (Omni-LLMs) have demonstrated impressive capabilities in holistic multi-modal perception, yet they consistently falter in complex scenarios requiring synergistic omni-modal reasoning. Beyond un…

Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

2026-01-31 · Zicheng Kong, Dehua Ma, Zhenbo Xu, Alven Yang 외 arxiv

Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly human labels, and provide opaque scalar…

Ola: Pushing the Frontiers of Omni-Modal Language Model

2025-02-06 · Zuyan Liu, Yuhao Dong, Jiahui Wang, Ziwei Liu 외

Recent advances in large language models, particularly following GPT-4o, have sparked increasing interest in developing omni-modal models capable of understanding more modalities. While some open-source alternatives have…

cross-modal alignmentLanguage ModelingLanguage Modelling