paper-with-me

Papers

Fashion Focus: Multi-modal Retrieval System for Video Commodity Localization in E-commerce

2021-02-09 · Yanhao Zhang, Qiang Wang, Pan Pan, Yun Zheng, Cheng Da, Siyang Sun, Yinghui Xu

Nowadays, live-stream and short video shopping in E-commerce have grown exponentially. However, the sellers are required to manually match images of the selling products to the timestamp of exhibition in the untrimmed video, resulting in a complicated process. To solve the problem, we present an innovative demonstration of multi-modal retrieval system called "Fashion Focus", which enables to exactly localize the product images in the online video as the focuses. Different modality contributes to the community localization, including visual content, linguistic features and interaction context are jointly investigated via presented multi-modal learning. Our system employs two procedures for analysis, including video content structuring and multi-modal retrieval, to automatically achieve accurate video-to-shop matching. Fashion Focus presents a unified framework that can orientate the consumers towards relevant product exhibitions during watching videos and help the sellers to effectively deliver the products over search and recommendation.

📄 PDF Abstract BibTeX arXiv:2102.04727

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalVideo-to-Shop

Similar Papers 제목 키워드 기반

UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation

2024-08-21 · Xiangyu Zhao, Yuehan Zhang, Wenlong Zhang, Xiao-Ming Wu

The fashion domain encompasses a variety of real-world multimodal tasks, including multimodal retrieval and multimodal generation. The rapid advancements in artificial intelligence generated content, particularly in tech…

Image GenerationImage RetrievalImage to textLanguage Modeling+4

FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and Captioning

2022-10-26 · Suvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha 외

Multimodal tasks in the fashion domain have significant potential for e-commerce, but involve challenging vision-and-language learning problems - e.g., retrieving a fashion item given a reference image plus text feedback…

Cross-Modal RetrievalDecoderFADImage Captioning+3

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

2026-05-21 · Haokun Wen, Xuemeng Song, Xinghao Xie, Xiaolin Chen 외 arxiv

Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is highly desired in practice. However, existing approaches focus on na…

Image Retrieval

FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training

2024-12-28 · Jiale Huang, Dehong Gao, Jinxia Zhang, Zechao Zhan 외

Large-scale Vision-Language Pre-training (VLP) has demonstrated remarkable success in the general domain. However, in the fashion domain, items are distinguished by fine-grained attributes like texture and material, whic…

AttributeImage ReconstructionRetrieval

Diverse-Intent Multi-Turn Fashion Image Retrieval

2026-07-22 · Mingqiang Tang, Haokun Wen, Meng Liu, Yupeng Hu 외 arxiv

Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every interaction follows the same attribute-editi…

Image Retrieval