paper-with-me

홈 › Papers

CurlingNet: Compositional Learning between Images and Text for Fashion IQ Data

2020-03-27 · Youngjae Yu, Seunghwan Lee, Yuncheol Choi, Gunhee Kim

We present an approach named CurlingNet that can measure the semantic distance of composition of image-text embedding. In order to learn an effective image-text composition for the data in the fashion domain, our model proposes two key components as follows. First, the Delivery makes the transition of a source image in an embedding space. Second, the Sweeping emphasizes query-related components of fashion images in the embedding space. We utilize a channel-wise gating mechanism to make it possible. Our single model outperforms previous state-of-the-art image-text composition models including TIRG and FiLM. We participate in the first fashion-IQ challenge in ICCV 2019, for which ensemble of our model achieves one of the best performances.

📄 PDF Abstract BibTeX arXiv:2003.12299

Code (1)

nashory/rtic-gcn-pytorch pytorch

Tasks

Image Retrieval

Similar Papers 제목 키워드 기반

FashionComposer: Compositional Fashion Image Generation

2024-12-18 · Sihui Ji, Yiyang Wang, Xi Chen, Xiaogang Xu 외

We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, parametric human model, garment image, an…

Image GenerationVirtual Try-on

Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval

2025-05-26 · Rong-Cheng Tu, Wenhao Sun, Hanzhe You, Yingjie Wang 외

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images given a compositional query, consisting of a reference image and a modifying text-without relying on annotated training data. Existing approaches…

Contrastive LearningImage RetrievalMultimodal ReasoningRetrieval+1

TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

2024-11-04 · Maitreya Patel, Abhiram Kusumba, Sheng Cheng, Changhoon Kim 외

Contrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the training data a significant factor in the eff…

Diversityimage-classificationImage ClassificationImage Retrieval+2

Learning Joint Visual Semantic Matching Embeddings for Language-guided Retrieval

2020-08-01 · ECCV 2020 8 · Yanbei Chen, Loris Bazzani

Interactive image retrieval is an emerging research topic with the objective of integrating inputs from multiple modalities as query for retrieval, e.g., textual feedback from users to guide, modify or refine image retri…

Image RetrievalRetrievalSpecificityText Retrieval

Image Search with Text Feedback by Additive Attention Compositional Learning

2022-03-08 · Yuxin Tian, Shawn Newsam, Kofi Boakye

Effective image retrieval with text feedback stands to impact a range of real-world applications, such as e-commerce. Given a source image and text feedback that describes the desired modifications to that image, the goa…

Image RetrievalRetrieval