paper-with-me

홈 › Papers

Bottleneck Tokens for Unified Multimodal Retrieval

2026-04-13 · Siyu Sun, Jing Ren, Zhaohe Liao, Dongxiao Mao, Xiangyuan Ren, Yiyi Zhang, Haohua Zhao, Weixiong Lin, Jiang Shaohua, Liqing Zhang, Yuchao Zheng arxiv

Adapting decoder-only multimodal large language models (MLLMs) for unified multimodal retrieval faces two structural gaps. First, existing methods rely on implicit pooling, which overloads the hidden state of a standard vocabulary token (e.g., <EOS>) as the sequence-level representation, a mechanism never designed for information aggregation. Second, contrastive fine-tuning specifies what the embedding should match but provides no token-level guidance on how information should be compressed into it. We address both gaps with two complementary components. Architecturally, we introduce Bottleneck Tokens (BToks), a small set of learnable tokens that serve as a fixed-capacity explicit pooling mechanism. For training, we propose Generative Information Condensation: a next-token prediction objective coupled with a Condensation Mask that severs the direct attention path from target tokens to query tokens. All predictive signals are thereby forced through the BToks, converting the generative loss into dense, token-level supervision for semantic compression. At inference time, only the input and BToks are processed in a single forward pass with negligible overhead over conventional last-token pooling. On MMEB-V2 (78 datasets, 3 modalities, 9 meta-tasks), our approach achieves state-of-the-art among 2B-scale methods under comparable data conditions, attaining an Overall score of 59.0 (+3.6 over VLM2Vec-V2) with substantial gains on semantically demanding tasks (e.g., +12.6 on Video-QA).

📄 PDF Abstract BibTeX arXiv:2604.11095

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Representation Forcing for Bottleneck-Free Unified Multimodal Models

2026-05-29 · Yuqing Wang, Zhijie Lin, Ceyuan Yang, Yang Zhao 외 arxiv

Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation, imposing a structural bottleneck. Naiv…

Image Generation

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

2026-09-15 · Guangyu Sun, Shlok Kumar Mishra, Wentao Bao, Robert Zhenheng Yang 외 arxiv

Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks g…

Representation LearningCross-Modal RetrievalImage Captioning

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression

2026-06-23 · Yijing Chen, Wenhui Tan, Xiaoyi Yu, Yuyue Wang 외 arxiv

Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information red…

Information Retrieval

Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking

2025-10-08 · Mitchell Keren Taraday, Shahaf Wagner, Chaim Baskin arxiv

Multimodal retrieval still leans on embedding-based models like CLIP for fast vector search over pre-computed image embeddings. Yet, unlike text retrieval, where joint-encoder rerankers are standard, comparable vision-la…

Text Retrieval

AMES: Approximate Multi-modal Enterprise Search via Late Interaction Retrieval

2026-03-13 · Tony Joseph, Carlos Pareja, David Lopes Pegna, Abhishek Singh arxiv

We present AMES (Approximate Multimodal Enterprise Search), a unified multimodal late interaction retrieval architecture which is backend agnostic. AMES demonstrates that fine-grained multimodal late interaction retrieva…

Cross-Modal Retrieval