paper-with-me

홈 › Papers

Closing the Modality Gap for Mixed Modality Search

2025-07-25 · Binxu Li, Yuhui Zhang, Xiaohan Wang, Weixin Liang, Ludwig Schmidt, Serena Yeung-Levy arxiv

Mixed modality search -- retrieving information across a heterogeneous corpus composed of images, texts, and multimodal documents -- is an important yet underexplored real-world application. In this work, we investigate how contrastive vision-language models, such as CLIP, perform on the mixed modality search task. Our analysis reveals a critical limitation: these models exhibit a pronounced modality gap in the embedding space, where image and text embeddings form distinct clusters, leading to intra-modal ranking bias and inter-modal fusion failure. To address this issue, we propose GR-CLIP, a lightweight post-hoc calibration method that removes the modality gap in CLIP's embedding space. Evaluated on MixBench -- the first benchmark specifically designed for mixed modality search -- GR-CLIP improves NDCG@10 by up to 26 percentage points over CLIP, surpasses recent vision-language generative embedding models by 4 percentage points, while using 75x less compute.

📄 PDF Abstract BibTeX arXiv:2507.19054

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RGB-T Tracking Based on Mixed Attention

2023-04-09 · Yang Luo, Xiqing Guo, Mingtao Dong, Jin Yu

RGB-T tracking involves the use of images from both visible and thermal modalities. The primary objective is to adaptively leverage the relatively dominant modality in varying conditions to achieve more robust tracking c…

Rgb-T Tracking

Multimodal Action Quality Assessment

2024-01-31 · Ling-An Zeng, Wei-Shi Zheng

Action quality assessment (AQA) is to assess how well an action is performed. Previous works perform modelling by only the use of visual information, ignoring audio information. We argue that although AQA is highly depen…

Action Quality AssessmentDecoderOptical Flow Estimation

From Cross-Modal to Mixed-Modal Visible-Infrared Re-Identification

2025-01-23 · Mahdi Alehdaghi, Rajarshi Bhattacharya, Pourya Shamsolmoali, Rafael M. O. Cruz 외

Visible-infrared person re-identification (VI-ReID) aims to match individuals across different camera modalities, a critical task in modern surveillance systems. While current VI-ReID methods focus on cross-modality matc…

Person Re-Identification

Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

2026-04-07 · Thibault Bañeras-Roux, Sergio Burdisso, Esaú Villatoro-Tello, Dairazalia Sánchez-Cortés 외 arxiv

Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR architectures connect a speech encoder to a large language model via a projec…

Speech RecognitionDomain Adaptation

Visible-Infrared Person Re-Identification via Patch-Mixed Cross-Modality Learning

2023-02-16 · Zhihao Qian, Yutian Lin, Bo Du

Visible-infrared person re-identification (VI-ReID) aims to retrieve images of the same pedestrian from different modalities, where the challenges lie in the significant modality discrepancy. To alleviate the modality ga…

Image GenerationPerson Re-IdentificationRepresentation LearningSemantic correspondence