paper-with-me

홈 › Papers

Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval

2024-05-23 · Young Kyun Jang, Donghyun Kim, Ser-Nam Lim

`Learning to hash'' is a practical solution for efficient retrieval, offering fast search speed and low storage cost. It is widely applied in various applications, such as image-text cross-modal search. In this paper, we explore the potential of enhancing the performance of learning to hash with the proliferation of powerful large pre-trained models, such as Vision-Language Pre-training (VLP) models. We introduce a novel method named Distillation for Cross-Modal Quantization (DCMQ), which leverages the rich semantic knowledge of VLP models to improve hash representation learning. Specifically, we use the VLP as a teacher' to distill knowledge into a `student' hashing model equipped with codebooks. This process involves the replacement of supervised labels, which are composed of multi-hot vectors and lack semantics, with the rich semantics of VLP. In the end, we apply a transformation termed Normalization with Paired Consistency (NPC) to achieve a discriminative target for distillation. Further, we introduce a new quantization method, Product Quantization with Gumbel (PQG) that promotes balanced codebook learning, thereby improving the retrieval performance. Extensive benchmark testing demonstrates that DCMQ consistently outperforms existing supervised cross-modal hashing approaches, showcasing its significant potential.

📄 PDF Abstract BibTeX arXiv:2405.14726

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalQuantizationRepresentation LearningRetrieval

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

GLaD: Geometric Latent Distillation for Vision-Language-Action Models

2025-12-10 · Minghao Guo, Meng Cao, Jiachen Tao, Rongtao Xu 외 arxiv

Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we introduce GLaD, a geometry-aware VLA fra…

Knowledge DistillationSpatial Reasoning

VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

2023-04-17 · Jing Liu, Sihan Chen, Xingjian He, Longteng Guo 외

In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation. Different from widely-studied vision-language pretraining models, VALOR jointly mo…

Audio captioningAudio-Video Question Answering (AVQA)Audio-Visual CaptioningAudio-visual Question Answering+16

SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment

2025-11-04 · Wenbo Lu arxiv

Vision-Language Pretraining (VLP) has achieved remarkable success across various downstream tasks, but such gains are largely driven by scaling up on training data. Yet, literature methods treat image-text pairs as isola…

Cross-Modal Retrieval

Image as a Foreign Language: BEiT Pretraining for Vision and Vision-Language Tasks

2023-01-01 · CVPR 2023 1 · Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck 외

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves excellent transfer performance on both vi…

Cross-Modal RetrievalImage Captioningimage-classificationImage Classification+10

Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

2022-08-22 · Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck 외

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both…

AllCross-Modal RetrievalImage Captioningimage-classification+13