paper-with-me

홈 › Papers

Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models

2017-11-17 · CVPR 2018 6 · Jiuxiang Gu, Jianfei Cai, Shafiq Joty, Li Niu, Gang Wang

Textual-visual cross-modal retrieval has been a hot research topic in both computer vision and natural language processing communities. Learning appropriate representations for multi-modal data is crucial for the cross-modal retrieval performance. Unlike existing image-text retrieval approaches that embed image-text pairs as single feature vectors in a common representational space, we propose to incorporate generative processes into the cross-modal feature embedding, through which we are able to learn not only the global abstract features but also the local grounded features. Extensive experiments show that our framework can well match images and sentences with complex content, and achieve the state-of-the-art cross-modal retrieval results on MSCOCO dataset.

📄 PDF Abstract BibTeX arXiv:1711.06420

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalImage-text RetrievalRetrievalText Retrieval

Similar Papers 제목 키워드 기반

DiffImaginE: Imagine to Verify Entity Types with Diffusion

2026-08-04 · Feng Zhang, Feiyu Han, Rongxin Yang, Yang Liu 외 arxiv

Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and-compare verifiers map each (span, type)…

IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval

2026-06-06 · Jiale Huang, Zixu Li, Zhiwei Chen, Zhiheng Fu 외 arxiv

Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing methods explore cross-modal correspondences, they often assume modified…

Image RetrievalVideo Retrieval

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning

2026-04-19 · Yian Li, Yang Jiao, Bin Zhu, Tianwen Qian 외 arxiv

Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, r…

multimodal generationSpatial Reasoning

Don't Just Listen, Use Your Imagination: Leveraging Visual Common Sense for Non-Visual Tasks

2015-02-21 · CVPR 2015 6 · Xiao Lin, Devi Parikh

Artificial agents today can answer factual questions. But they fall short on questions that require common sense reasoning. Perhaps this is because most existing common sense databases rely on text to learn and represent…

Common Sense Reasoning

Enhancing Zero-shot Commonsense Reasoning by Integrating Visual Knowledge via Machine Imagination

2026-03-05 · Hyuntae Park, Yeachan Kim, SangKeun Lee arxiv

Recent advancements in zero-shot commonsense reasoning have empowered Pre-trained Language Models (PLMs) to acquire extensive commonsense knowledge without requiring task-specific fine-tuning. Despite this progress, thes…