paper-with-me

Papers

Multi-Modal Fashion Product Retrieval

2017-04-01 · WS 2017 4 · Antonio Rubio Romano, LongLong Yu, Edgar Simo-Serra, Francesc Moreno-Noguer

Finding a product in the fashion world can be a daunting task. Everyday, e-commerce sites are updating with thousands of images and their associated metadata (textual information), deepening the problem. In this paper, we leverage both the images and textual metadata and propose a joint multi-modal embedding that maps both the text and images into a common latent space. Distances in the latent space correspond to similarity between products, allowing us to effectively perform retrieval in this latent space. We compare against existing approaches and show significant improvements in retrieval tasks on a large-scale e-commerce dataset.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Similar Papers 제목 키워드 기반

Fashion Focus: Multi-modal Retrieval System for Video Commodity Localization in E-commerce

2021-02-09 · Yanhao Zhang, Qiang Wang, Pan Pan, Yun Zheng 외

Nowadays, live-stream and short video shopping in E-commerce have grown exponentially. However, the sellers are required to manually match images of the selling products to the timestamp of exhibition in the untrimmed vi…

RetrievalVideo-to-Shop

FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data

2026-04-11 · Peng Yuan, Bingyin Mei, Hui Zhang arxiv

Composed Image Retrieval (CIR) retrieves target images using a reference image paired with modification text. Despite rapid advances, all existing methods and datasets operate at the image level -- a single reference ima…

Image Retrieval

Progressive Learning for Image Retrieval with Hybrid-Modality Queries

2022-04-24 · Yida Zhao, Yuqing Song, Qin Jin

Image retrieval with hybrid-modality queries, also known as composing text and image for image retrieval (CTI-IR), is a retrieval task where the search intention is expressed in a more complex query format, involving bot…

Image RetrievalImage-text RetrievalRetrievalText Retrieval

UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation

2024-08-21 · Xiangyu Zhao, Yuehan Zhang, Wenlong Zhang, Xiao-Ming Wu

The fashion domain encompasses a variety of real-world multimodal tasks, including multimodal retrieval and multimodal generation. The rapid advancements in artificial intelligence generated content, particularly in tech…

Image GenerationImage RetrievalImage to textLanguage Modeling+4

FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and Captioning

2022-10-26 · Suvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha 외

Multimodal tasks in the fashion domain have significant potential for e-commerce, but involve challenging vision-and-language learning problems - e.g., retrieving a fashion item given a reference image plus text feedback…

Cross-Modal RetrievalDecoderFADImage Captioning+3