Semantic-enhanced Modality-asymmetric Retrieval for Online E-commerce Search
Semantic retrieval, which retrieves semantically matched items given a textual query, has been an essential component to enhance system effectiveness in e-commerce search. In this paper, we study the multimodal retrieval problem, where the visual information (e.g, image) of item is leveraged as supplementary of textual information to enrich item representation and further improve retrieval performance. Though learning from cross-modality data has been studied extensively in tasks such as visual question answering or media summarization, multimodal retrieval remains a non-trivial and unsolved problem especially in the asymmetric scenario where the query is unimodal while the item is multimodal. In this paper, we propose a novel model named SMAR, which stands for Semantic-enhanced Modality-Asymmetric Retrieval, to tackle the problem of modality fusion and alignment in this kind of asymmetric scenario. Extensive experimental results on an industrial dataset show that the proposed model outperforms baseline models significantly in retrieval accuracy. We have open sourced our industrial dataset for the sake of reproducibility and future research works.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringRetrievalSemantic RetrievalVisual Question AnsweringSimilar Papers 제목 키워드 기반
Task-adaptive Asymmetric Deep Cross-modal Hashing
Supervised cross-modal hashing aims to embed the semantic correlations of heterogeneous modality data into the binary hash codes with discriminative semantic labels. Because of its advantages on retrieval and storage eff…
Cross-Modal RetrievalRetrievalAdvancing Drug Discovery with Enhanced Chemical Understanding via Asymmetric Contrastive Multimodal Learning
The versatility of multimodal deep learning holds tremendous promise for advancing scientific research and practical applications. As this field continues to evolve, the collective power of cross-modal analysis promises …
Contrastive LearningDrug DiscoveryMolecular Property PredictionMultimodal Deep Learning+3Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval
Video Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often s…
cross-modal alignmentMoment RetrievalRetrievalSentenceREAD: Retrieval-Enhanced Asymmetric Diffusion for Motion Planning
This paper proposes Retrieval-Enhanced Asymmetric Diffusion (READ) for image-based robot motion planning. Given an image of the scene READ retrieves an initial motion from a database of image-motion pairs and uses a …
Motion PlanningRetrievalShrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision Alignment
Decoding visual features from EEG signals is a central challenge in neuroscience, with cross-modal alignment as the dominant approach. We argue that the relationship between visual and brain modalities is fundamentally a…
Image Retrieval