paper-with-me

홈 › Papers

From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model

2025-08-01 · Yeong-Joon Ju, Seong-Whan Lee arxiv

Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrastive pre-training, while traditional hard negative mining methods suffer from severe false negative contamination. In this paper, we propose a highly data-efficient framework that bypasses extensive pre-training to build a robust multimodal representation space. We first introduce a hierarchical embedding prompt that provides strong latent conditioning. By explicitly anchoring task definitions at the system level, this prompting strategy effectively bridges the modality gap and unlocks powerful zero-shot embedding capabilities. Building upon this latent conditioning, we present Self-aware Hard Negative Sampling (SaHa). Unlike conventional candidate-space mining, SaHa shifts the mechanism to the query-space by mapping retrieved candidates back to their owner queries to rigorously filter out semantic false negatives. Furthermore, our method constructs mutually hard clusters, maximizing intra-task discrimination and batch efficiency without redundant forward passes. Extensive experiments demonstrate that our unified approach achieves highly competitive fine-tuning performance on the Massive Multimodal Embedding Benchmark using only a fraction of standard training data.

📄 PDF Abstract BibTeX arXiv:2508.00955

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Omni-Interactive Universal Embedder

2026-08-27 · Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon 외 arxiv

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, e…

Representation Learning

E3: Ensemble of Expert Embedders for Adapting Synthetic Image Detectors to New Generators Using Limited Data

2024-04-12 · Aref Azizpour, Tai D. Nguyen, Manil Shrestha, Kaidi Xu 외

As generative AI progresses rapidly, new synthetic image generators continue to emerge at a swift pace. Traditional detection methods face two main challenges in adapting to these generators: the forensic traces of synth…

Continual LearningSynthetic Image DetectionTransfer Learning

Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation

2025-11-04 · Wongyu Kim, Hochang Lee, Sanghak Lee, Yoonsung Kim 외 arxiv

Query augmentation makes queries more meaningful by appending further information to the queries to find relevant documents. Current studies have proposed Large Language Model (LLM)-based embedders, which learn represent…

Unlocking Innate Computing Abilities in Electric Grids

2025-05-15 · Yubo Song, Subham Sahoo

High energy consumption of artificial intelligence has gained momentum worldwide, which necessitates major investments on expanding efficient and carbon-neutral generation and data center infrastructure in electric power…

VIRTUE: Visual-Interactive Text-Image Universal Embedder

2025-10-01 · Wei-Yao Wang, Kazuya Tateishi, Qiyu Wu, Shusuke Takahashi 외 arxiv

Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embedding models with instruction-following …

Representation Learning