Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models
Textual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval performance. Unlike existing image-text retrieval approaches that embed image-text pairs as single feature vectors in a common representational space, we propose to incorporate generative processes into the cross-modal feature embedding, through which we are able to learn not only the global abstract features but also the local grounded features. Extensive experiments show that our framework can well match images and sentences with complex content, and achieve the state-of-the-art cross-modal retrieval results on MSCOCO dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Modal RetrievalImage-text RetrievalRetrievalText RetrievalSimilar Papers 제목 키워드 기반
DiffImaginE: Imagine to Verify Entity Types with Diffusion
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type)…
IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval
Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing methods explore cross-modal correspondences, they often assume modified…
Image RetrievalVideo RetrievalSpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, r…
multimodal generationSpatial ReasoningDon't Just Listen, Use Your Imagination: Leveraging Visual Common Sense for Non-Visual Tasks
Artificial agents today can answer factual questions. But they fall short on questions that require common sense reasoning. Perhaps this is because most existing common sense databases rely on text to learn and represent…
Common Sense ReasoningEnhancing Zero-shot Commonsense Reasoning by Integrating Visual Knowledge via Machine Imagination
Recent advancements in zero-shot commonsense reasoning have empowered Pre-trained Language Models (PLMs) to acquire extensive commonsense knowledge without requiring task-specific fine-tuning. Despite this progress, thes…