paper-with-me

홈 › Papers

One-Shot Item Search with Multimodal Data

2018-11-27 · Jonghwa Yim, Junghun James Kim, Daekyu Shin

In the task of near similar image search, features from Deep Neural Network is often used to compare images and measure similarity. In the past, we only focused visual search in image dataset without text data. However, since deep neural network emerged, the performance of visual search becomes high enough to apply it in many industries from 3D data to multimodal data. Compared to the needs of multimodal search, there has not been sufficient researches. In this paper, we present a method of near similar search with image and text multimodal dataset. Earlier time, similar image search, especially when searching shopping items, treated image and text separately to search similar items and reorder the results. This regards two tasks of image search and text matching as two different tasks. Our method, however, explore the vast data to compute k-nearest neighbors using both image and text. In our experiment of similar item search, our system using multimodal data shows better performance than single data while it only increases minute computing time. For the experiment, we collected more than 15 million of accessory and six million of digital product items from online shopping websites, in which the product item comprises item images, titles, categories, and descriptions. Then we compare the performance of multimodal searching to single space searching in these datasets.

📄 PDF Abstract BibTeX arXiv:1811.10969

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalText Matching

Similar Papers 제목 키워드 기반

Zero-Shot Recommendations with Pre-Trained Large Language Models for Multimodal Nudging

2023-09-02 · Rachel M. Harrison, Anton Dereventsov, Anton Bibin

We present a method for zero-shot recommendation of multimodal non-stationary content that leverages recent advancements in the field of generative AI. We propose rendering inputs of different modalities as textual descr…

Zero-Shot Next-Item Recommendation using Large Pretrained Language Models

2023-04-06 · Lei Wang, Ee-Peng Lim

Large language models (LLMs) have achieved impressive zero-shot performance in various natural language processing (NLP) tasks, demonstrating their capabilities for inference without training examples. Despite their succ…

Sequential Recommendation

Learning Visuo-Haptic Skewering Strategies for Robot-Assisted Feeding

2022-11-26 · Priya Sundaresan, Suneel Belkhale, Dorsa Sadigh

Acquiring food items with a fork poses an immense challenge to a robot-assisted feeding system, due to the wide range of material properties and visual appearances present across food groups. Deformable foods necessitate…

Diversity

Sens-VisualNews: A Benchmark Dataset for Sensational Image Detection

2026-05-11 · Andreas Goulas, Damianos Galanopoulos, Evlampios Apostolidis, Vasileios Mezaris arxiv

The detection of sensational content in media items can be a critical filtering mechanism for identifying check-worthy content and flagging potential disinformation, since such content triggers physiological arousal that…

Semantic-enhanced Modality-asymmetric Retrieval for Online E-commerce Search

2025-06-25 · Zhigong Zhou, Ning Ding, Xiaochuan Fan, Yue Shang 외

Semantic retrieval, which retrieves semantically matched items given a textual query, has been an essential component to enhance system effectiveness in e-commerce search. In this paper, we study the multimodal retrieval…

Question AnsweringRetrievalSemantic RetrievalVisual Question Answering