paper-with-me

홈 › Papers

Images Don't Lie: Transferring Deep Visual Semantic Features to Large-Scale Multimodal Learning to Rank

2015-11-20 · Corey Lynch, Kamelia Aryafar, Josh Attenberg

Search is at the heart of modern e-commerce. As a result, the task of ranking search results automatically (learning to rank) is a multibillion dollar machine learning problem. Traditional models optimize over a few hand-constructed features based on the item's text. In this paper, we introduce a multimodal learning to rank model that combines these traditional features with visual semantic features transferred from a deep convolutional neural network. In a large scale experiment using data from the online marketplace Etsy, we verify that moving to a multimodal representation significantly improves ranking quality. We show how image features can capture fine-grained style information not available in a text-only representation. In addition, we show concrete examples of how image information can successfully disentangle pairs of highly different items that are ranked similarly by a text-only model.

📄 PDF Abstract BibTeX arXiv:1511.06746

Code (0)

등록된 구현이 없습니다.

Tasks

Learning-To-Rank

Similar Papers 제목 키워드 기반

VT-CLIP: Enhancing Vision-Language Models with Visual-guided Texts

2021-12-04 · Longtian Qiu, Renrui Zhang, Ziyu Guo, Ziyao Zeng 외

Contrastive Language-Image Pre-training (CLIP) has drawn increasing attention recently for its transferable visual representation learning. However, due to the semantic gap within datasets, CLIP's pre-trained image-text …

Language ModellingRepresentation LearningZero-Shot Learning

Geometric Visual Similarity Learning in 3D Medical Image Self-supervised Pre-training

2023-03-02 · CVPR 2023 1 · Yuting He, Guanyu Yang, Rongjun Ge, Yang Chen 외

Learning inter-image similarity is crucial for 3D medical images self-supervised pre-training, due to their sharing of numerous same semantic regions. However, the lack of the semantic prior in metrics and the semantic-i…

Geometric MatchingRepresentation Learning

Multi-scale Semantic Prior Features Guided Deep Neural Network for Urban Street-view Image

2024-05-17 · Jianshun Zeng, Wang Li, Yanjie Lv, Shuai Gao 외

Street-view image has been widely applied as a crucial mobile mapping data source. The inpainting of street-view images is a critical step for street-view image processing, not only for the privacy protection, but also f…

DecoderImage Inpainting

SEGA: Semantic Guided Attention on Visual Prototype for Few-Shot Learning

2021-11-08 · Fengyuan Yang, Ruiping Wang, Xilin Chen

Teaching machines to recognize a new category based on few training samples especially only one remains challenging owing to the incomprehensive understanding of the novel category caused by the lack of data. However, hu…

feature selectionFew-Shot Learning

Epsilon: Exploring Comprehensive Visual-Semantic Projection for Multi-Label Zero-Shot Learning

2024-08-22 · Ziming Liu, Jingcai Guo, Song Guo, Xiaocheng Lu

This paper investigates a challenging problem of zero-shot learning in the multi-label scenario (MLZSL), wherein the model is trained to recognize multiple unseen classes within a sample (e.g., an image) based on seen cl…

Multi-label zero-shot learningZero-Shot Learning