paper-with-me

홈 › Papers

Dialog-based Interactive Image Retrieval

2018-05-01 · NeurIPS 2018 12 · Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, Rogerio Schmidt Feris

Existing methods for interactive image retrieval have demonstrated the merit of integrating user feedback, improving retrieval results. However, most current systems rely on restricted forms of user feedback, such as binary relevance responses, or feedback based on a fixed set of relative attributes, which limits their impact. In this paper, we introduce a new approach to interactive image search that enables users to provide feedback via natural language, allowing for more natural and effective interaction. We formulate the task of dialog-based interactive image retrieval as a reinforcement learning problem, and reward the dialog system for improving the rank of the target image during each dialog turn. To mitigate the cumbersome and costly process of collecting human-machine conversations as the dialog system learns, we train our system with a user simulator, which is itself trained to describe the differences between target and candidate images. The efficacy of our approach is demonstrated in a footwear retrieval application. Experiments on both simulated and real-world data show that 1) our proposed learning framework achieves better accuracy than other supervised and reinforcement learning baselines and 2) user feedback based on natural language rather than pre-specified attributes leads to more effective retrieval results, and a more natural and expressive communication interface.

📄 PDF Abstract BibTeX arXiv:1805.00145

Code (1)

XiaoxiaoGuo/fashion-retrieval 공식 구현 pytorch

Tasks

Image Retrievalreinforcement-learningReinforcement LearningReinforcement Learning (RL)RetrievalVisual Dialog

Similar Papers 제목 키워드 기반

DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval

2025-11-18 · Zongwei Zhen, Biqing Zeng arxiv

This paper addresses the task of interactive, conversational text-to-image retrieval. Our DIR-TIR framework progressively refines the target image search through two specialized modules: the Dialog Refiner Module and the…

Image Retrieval

Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach

2024-06-05 · Saehyung Lee, Sangwon Yu, Junsung Park, Jihun Yi 외

In this paper, we primarily address the issue of dialogue-form context query within the interactive text-to-image retrieval task. Our methodology, PlugIR, actively utilizes the general instruction-following capability of…

Image RetrievalInstruction FollowingRetrieval

Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval with Partial Query

2021-03-02 · ICCV 2021 10 · Guanyu Cai, Jun Zhang, Xinyang Jiang, Yifei Gong 외

Text-based image retrieval has seen considerable progress in recent years. However, the performance of existing methods suffers in real life since the user is likely to provide an incomplete description of an image, whic…

Cross-Modal RetrievalImage RetrievalRetrieval

Fashion IQ: A New Dataset Towards Retrieving Images by Natural Language Feedback

2019-05-30 · CVPR 2021 1 · Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah 외

Conversational interfaces for the detail-oriented retail fashion domain are more natural, expressive, and user friendly than classical keyword-based search interfaces. In this paper, we introduce the Fashion IQ dataset t…

AttributeImage RetrievalRetrieval

Chat-based Person Retrieval via Dialogue-Refined Cross-Modal Alignment

2025-01-01 · CVPR 2025 1 · Yang Bai, Yucheng Ji, Min Cao, Jinqiao Wang 외

Traditional text-based person retrieval (TPR) relies on a single-shot text as query to retrieve the target person, assuming that the query completely captures the user's search intent. However, in real-world scenario…

Attributecross-modal alignmentData AugmentationPerson Retrieval+5