ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval
Composed image retrieval (CIR) is the task of retrieving a target image specified by a query image and a relative text that describes a semantic modification to the query image. Existing methods in CIR struggle to accurately represent the image and the text modification, resulting in subpar performance. To address this limitation, we introduce a CIR framework, ConText-CIR, trained with a Text Concept-Consistency loss that encourages the representations of noun phrases in the text modification to better attend to the relevant parts of the query image. To support training with this loss function, we also propose a synthetic data generation pipeline that creates training data from existing CIR datasets or unlabeled images. We show that these components together enable stronger performance on CIR tasks, setting a new state-of-the-art in composed image retrieval in both the supervised and zero-shot settings on multiple benchmark datasets, including CIRR and CIRCO. Source code, model checkpoints, and our new datasets are available at https://github.com/mvrl/ConText-CIR.
Code (1)
Tasks
Image RetrievalRetrievalSynthetic Data GenerationSimilar Papers 제목 키워드 기반
NEUCORE: Neural Concept Reasoning for Composed Image Retrieval
Composed image retrieval which combines a reference image and a text modifier to identify the desired target image is a challenging task, and requires the model to comprehend both vision and language modalities and their…
Concept AlignmentImage RetrievalMultiple Instance LearningRetrieval+1IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval
Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing methods explore cross-modal correspondences, they often assume modified…
Image RetrievalVideo RetrievalKnowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval
We study the zero-shot Composed Image Retrieval (ZS-CIR) task, which is to retrieve the target image given a reference image and a description without training on the triplet datasets. Previous works generate pseudo-word…
AttributeImage RetrievalRetrievalTriplet+1Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
Text-to-image diffusion models (T2I) use a latent representation of a text prompt to guide the image generation process. However, the process by which the encoder produces the text representation is unknown. We propose t…
Image GenerationRetrievalHINT: Composed Image Retrieval with Dual-path Compositional Contextualized Network
Composed Image Retrieval (CIR) is a challenging image retrieval paradigm. It aims to retrieve target images from large-scale image databases that are consistent with the modification semantics, based on a multimodal quer…
Image Retrieval