paper-with-me

홈 › Papers

FashionNTM: Multi-turn Fashion Image Retrieval via Cascaded Memory

2023-08-20 · ICCV 2023 1 · Anwesan Pal, Sahil Wadhwa, Ayush Jaiswal, Xu Zhang, Yue Wu, Rakesh Chada, Pradeep Natarajan, Henrik I. Christensen

Multi-turn textual feedback-based fashion image retrieval focuses on a real-world setting, where users can iteratively provide information to refine retrieval results until they find an item that fits all their requirements. In this work, we present a novel memory-based method, called FashionNTM, for such a multi-turn system. Our framework incorporates a new Cascaded Memory Neural Turing Machine (CM-NTM) approach for implicit state management, thereby learning to integrate information across all past turns to retrieve new images, for a given turn. Unlike vanilla Neural Turing Machine (NTM), our CM-NTM operates on multiple inputs, which interact with their respective memories via individual read and write heads, to learn complex relationships. Extensive evaluation results show that our proposed method outperforms the previous state-of-the-art algorithm by 50.5%, on Multi-turn FashionIQ -- the only existing multi-turn fashion dataset currently, in addition to having a relative improvement of 12.6% on Multi-turn Shoes -- an extension of the single-turn Shoes dataset that we created in this work. Further analysis of the model in a real-world interactive setting demonstrates two important capabilities of our model -- memory retention across turns, and agnosticity to turn order for non-contradictory feedback. Finally, user study results show that images retrieved by FashionNTM were favored by 83.1% over other multi-turn models. Project page: https://sites.google.com/eng.ucsd.edu/fashionntm

📄 PDF Abstract BibTeX arXiv:2308.10170

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalRetrieval

Methods 이 논문이 사용한 방법론

Tanh Activation 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Location-based Attention 설명 없음
Sigmoid Activation 설명 없음
Content-based Attention Content-based attention is an attention mechanism based on cosine similarity: $$f_{att}\left(\textbf{h}_{i}, \textbf{s}\_{j}\right) =…
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…
Neural Turing Machine A Neural Turing Machine is a working memory neural network model. It couples a neural network architecture with external memory resources. The whole architecture is…

Similar Papers 제목 키워드 기반

Conversational Fashion Image Retrieval via Multiturn Natural Language Feedback

2021-06-08 · Yifei Yuan, Wai Lam

We study the task of conversational fashion image retrieval via multiturn natural language feedback. Most previous studies are based on single-turn settings. Existing models on multiturn conversational fashion image retr…

AttributeImage RetrievalRetrieval

Diverse-Intent Multi-Turn Fashion Image Retrieval

2026-07-22 · Mingqiang Tang, Haokun Wen, Meng Liu, Yupeng Hu 외 arxiv

Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every interaction follows the same attribute-editi…

Image Retrieval

CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains

2026-05-26 · Tomohisa Takeda, Yu-Chieh Lin, Yuji Nozawa, Youyang Ng 외 arxiv

Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-historyconsistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, an…

Image Retrieval

UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation

2024-08-21 · Xiangyu Zhao, Yuehan Zhang, Wenlong Zhang, Xiao-Ming Wu

The fashion domain encompasses a variety of real-world multimodal tasks, including multimodal retrieval and multimodal generation. The rapid advancements in artificial intelligence generated content, particularly in tech…

Image GenerationImage RetrievalImage to textLanguage Modeling+4

FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and Captioning

2022-10-26 · Suvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha 외

Multimodal tasks in the fashion domain have significant potential for e-commerce, but involve challenging vision-and-language learning problems - e.g., retrieving a fashion item given a reference image plus text feedback…

Cross-Modal RetrievalDecoderFADImage Captioning+3