paper-with-me

홈 › Papers

ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval

2025-05-27 · CVPR 2025 1 · Eric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou, Nathan Jacobs

Composed image retrieval (CIR) is the task of retrieving a target image specified by a query image and a relative text that describes a semantic modification to the query image. Existing methods in CIR struggle to accurately represent the image and the text modification, resulting in subpar performance. To address this limitation, we introduce a CIR framework, ConText-CIR, trained with a Text Concept-Consistency loss that encourages the representations of noun phrases in the text modification to better attend to the relevant parts of the query image. To support training with this loss function, we also propose a synthetic data generation pipeline that creates training data from existing CIR datasets or unlabeled images. We show that these components together enable stronger performance on CIR tasks, setting a new state-of-the-art in composed image retrieval in both the supervised and zero-shot settings on multiple benchmark datasets, including CIRR and CIRCO. Source code, model checkpoints, and our new datasets are available at https://github.com/mvrl/ConText-CIR.

📄 PDF Abstract BibTeX arXiv:2505.20764

Code (1)

mvrl/context-cir 공식 구현 pytorch

Tasks

Image RetrievalRetrievalSynthetic Data Generation

Similar Papers 제목 키워드 기반

NEUCORE: Neural Concept Reasoning for Composed Image Retrieval

2023-10-02 · Shu Zhao, Huijuan Xu

Composed image retrieval which combines a reference image and a text modifier to identify the desired target image is a challenging task, and requires the model to comprehend both vision and language modalities and their…

Concept AlignmentImage RetrievalMultiple Instance LearningRetrieval+1

IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval

2026-06-06 · Jiale Huang, Zixu Li, Zhiwei Chen, Zhiheng Fu 외 arxiv

Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing methods explore cross-modal correspondences, they often assume modified…

Image RetrievalVideo Retrieval

Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval

2024-03-24 · CVPR 2024 1 · Yucheng Suo, Fan Ma, Linchao Zhu, Yi Yang

We study the zero-shot Composed Image Retrieval (ZS-CIR) task, which is to retrieve the target image given a reference image and a description without training on the triplet datasets. Previous works generate pseudo-word…

AttributeImage RetrievalRetrievalTriplet+1

Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines

2024-03-09 · Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad 외

Text-to-image diffusion models (T2I) use a latent representation of a text prompt to guide the image generation process. However, the process by which the encoder produces the text representation is unknown. We propose t…

Image GenerationRetrieval

HINT: Composed Image Retrieval with Dual-path Compositional Contextualized Network

2026-03-27 · Mingyu Zhang, Zixu Li, Zhiwei Chen, Zhiheng Fu 외 arxiv

Composed Image Retrieval (CIR) is a challenging image retrieval paradigm. It aims to retrieve target images from large-scale image databases that are consistent with the modification semantics, based on a multimodal quer…

Image Retrieval