paper-with-me

홈 › Papers

Language-only Training of Zero-shot Composed Image Retrieval

2024-01-01 · CVPR 2024 1 · Geonmo Gu, Sanghyuk Chun, Wonjae Kim, Yoohoon Kang, Sangdoo Yun

Composed image retrieval (CIR) task takes a composed query of image and text aiming to search relative images for both conditions. Conventional CIR approaches need a training dataset composed of triplets of query image query text and target image which is very expensive to collect. Several recent works have worked on the zero-shot (ZS) CIR paradigm to tackle the issue without using pre-collected triplets. However the existing ZS-CIR methods show limited backbone scalability and generalizability due to the lack of diversity of the input texts during training. We propose a novel CIR framework only using language for its training. Our LinCIR (Language-only training for CIR) can be trained only with text datasets by a novel self-supervision named self-masking projection (SMP). We project the text latent embedding to the token embedding space and construct a new text by replacing the keyword tokens of the original text. Then we let the new and original texts have the same latent embedding vector. With this simple strategy LinCIR is surprisingly efficient and highly effective; LinCIR with CLIP ViT-G backbone is trained in 48 minutes and shows the best ZS-CIR performances on four different CIR benchmarks CIRCO GeneCIS FashionIQ and CIRR even outperforming supervised method on FashionIQ. Code is available at https://github.com/navervision/lincir

📄 PDF Abstract BibTeX

Code (1)

navervision/lincir 공식 구현 pytorch

Tasks

Image RetrievalRetrieval

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Language-only Efficient Training of Zero-shot Composed Image Retrieval

2023-12-04 · Geonmo Gu, Sanghyuk Chun, Wonjae Kim, Yoohoon Kang 외

Composed image retrieval (CIR) task takes a composed query of image and text, aiming to search relative images for both conditions. Conventional CIR approaches need a training dataset composed of triplets of query image,…

Image RetrievalRetrievalZero-Shot Composed Image Retrieval (ZS-CIR)

Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data

2025-04-01 · Yiqun Duan, Sameera Ramasinghe, Stephen Gould, Ajanthan Thalaiyasingam

Composed Image Retrieval (CIR) is the task of retrieving images matching a reference image augmented with a text, where the text describes changes to the reference image in natural language. Traditionally, models designe…

Image RetrievalRetrievalTriplet

HyCIR: Boosting Zero-Shot Composed Image Retrieval with Synthetic Labels

2024-07-08 · Yingying Jiang, Hanchao Jia, Xiaobing Wang, Peng Hao

Composed Image Retrieval (CIR) aims to retrieve images based on a query image with text. Current Zero-Shot CIR (ZS-CIR) methods try to solve CIR tasks without using expensive triplet-labeled training datasets. However, t…

Contrastive LearningImage RetrievalImage to textLanguage Modelling+4

MoTaDual: Modality-Task Dual Alignment for Enhanced Zero-shot Composed Image Retrieval

2024-10-31 · Haiwen Li, Fei Su, Zhicheng Zhao

Composed Image Retrieval (CIR) is a challenging vision-language task, utilizing bi-modal (image+text) queries to retrieve target images. Despite the impressive performance of supervised CIR, the dependence on costly, man…

Image RetrievalPrompt LearningRetrievalTriplet+1

Billions of Parameters Are Worth More Than In-domain Training Data: A case study in the Legal Case Entailment Task

2022-05-30 · Guilherme Moraes Rosa, Luiz Bonifacio, Vitor Jeronymo, Hugo Abonizio 외

Recent work has shown that language models scaled to billions of parameters, such as GPT-3, perform remarkably well in zero-shot and few-shot scenarios. In this work, we experiment with zero-shot models in the legal case…

Language Modelling