paper-with-me

홈 › Papers

WAD-CMSN: Wasserstein Distance based Cross-Modal Semantic Network for Zero-Shot Sketch-Based Image Retrieval

2022-02-11 · Guanglong Xu, Zhensheng Hu, Jia Cai

Zero-shot sketch-based image retrieval (ZSSBIR), as a popular studied branch of computer vision, attracts wide attention recently. Unlike sketch-based image retrieval (SBIR), the main aim of ZSSBIR is to retrieve natural images given free hand-drawn sketches that may not appear during training. Previous approaches used semantic aligned sketch-image pairs or utilized memory expensive fusion layer for projecting the visual information to a low dimensional subspace, which ignores the significant heterogeneous cross-domain discrepancy between highly abstract sketch and relevant image. This may yield poor performance in the training phase. To tackle this issue and overcome this drawback, we propose a Wasserstein distance based cross-modal semantic network (WAD-CMSN) for ZSSBIR. Specifically, it first projects the visual information of each branch (sketch, image) to a common low dimensional semantic subspace via Wasserstein distance in an adversarial training manner. Furthermore, identity matching loss is employed to select useful features, which can not only capture complete semantic knowledge, but also alleviate the over-fitting phenomenon caused by the WAD-CMSN model. Experimental results on the challenging Sketchy (Extended) and TU-Berlin (Extended) datasets indicate the effectiveness of the proposed WAD-CMSN model over several competitors.

📄 PDF Abstract BibTeX arXiv:2202.05465

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalRetrievalSketch-Based Image Retrieval

Similar Papers 제목 키워드 기반

Vision-Based Perception for Autonomous Vehicles in Off-Road Environment Using Deep Learning

2025-09-20 · Nelson Alves Ferreira Neto arxiv

Low-latency intelligent systems are required for autonomous driving on non-uniform terrain in open-pit mines and developing countries. This work proposes a perception system for autonomous vehicles on unpaved roads and o…

Real-Time Semantic SegmentationAutonomous VehiclesAutonomous Driving

Low-latency Perception in Off-Road Dynamical Low Visibility Environments

2020-12-23 · Nelson Alves, Marco Ruiz, Marco Reis, Tiago Cajahyba 외

This work proposes a perception system for autonomous vehicles and advanced driver assistance specialized on unpaved roads and off-road environments. In this research, the authors have investigated the behavior of Deep L…

Autonomous VehiclesSegmentationSemantic Segmentation

RecGOAT: Graph Optimal Adaptive Transport for LLM-Enhanced Multimodal Recommendation with Dual Semantic Alignment

2026-01-31 · Yuecheng Li, Hengwei Ju, Zeyu Song, Wei Yang 외 arxiv

Integrating large language model (LLM) representations into multimodal recommendation has shown promise, yet a fundamental challenge remains largely overlooked: the semantic heterogeneity between generative LM representa…

Multimodal RecommendationRecommendation SystemsContrastive Learning

Wasserstein distances for evaluating cross-lingual embeddings

2019-10-24 · Georgios Balikas, Ioannis Partalas

Word embeddings are high dimensional vector representations of words that capture their semantic similarity in the vector space. There exist several algorithms for learning such embeddings both for a single language as w…

Cross-Lingual Document ClassificationDocument ClassificationRetrievalSemantic Similarity+2

Joint Wasserstein Autoencoders for Aligning Multimodal Embeddings

2019-09-14 · Shweta Mahajan, Teresa Botschen, Iryna Gurevych, Stefan Roth

One of the key challenges in learning joint embeddings of multiple modalities, e.g. of images and text, is to ensure coherent cross-modal semantics that generalize across datasets. We propose to address this through join…

Cross-Modal RetrievalRetrieval