paper-with-me

홈 › Papers

Beyond the Deep Metric Learning: Enhance the Cross-Modal Matching with Adversarial Discriminative Domain Regularization

2020-10-23 · Li Ren, Kai Li, Liqiang Wang, Kien Hua

Matching information across image and text modalities is a fundamental challenge for many applications that involve both vision and natural language processing. The objective is to find efficient similarity metrics to compare the similarity between visual and textual information. Existing approaches mainly match the local visual objects and the sentence words in a shared space with attention mechanisms. The matching performance is still limited because the similarity computation is based on simple comparisons of the matching features, ignoring the characteristics of their distribution in the data. In this paper, we address this limitation with an efficient learning objective that considers the discriminative feature distributions between the visual objects and sentence words. Specifically, we propose a novel Adversarial Discriminative Domain Regularization (ADDR) learning framework, beyond the paradigm metric learning objective, to construct a set of discriminative data domains within each image-text pairs. Our approach can generally improve the learning efficiency and the performance of existing metrics learning frameworks by regulating the distribution of the hidden space between the matching pairs. The experimental results show that this new approach significantly improves the overall performance of several popular cross-modal matching techniques (SCAN, VSRN, BFAN) on the MS-COCO and Flickr30K benchmarks.

📄 PDF Abstract BibTeX arXiv:2010.12126

Code (0)

등록된 구현이 없습니다.

Tasks

Metric LearningSentence

Similar Papers 제목 키워드 기반

Cross Spatial Temporal Fusion Attention for Remote Sensing Object Detection via Image Feature Matching

2025-07-25 · Abu Sadat Mohammad Salehin Amit, Xiaoli Zhang, Md Masum Billa Shagar, Zhaojun Liu 외 arxiv

Effectively describing features for cross-modal remote sensing image matching remains a challenging task due to the significant geometric and radiometric differences between multimodal images. Existing methods primarily …

Computational EfficiencyObject DetectionImage Matching

MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training

2025-01-13 · Xingyi He, Hao Yu, Sida Peng, Dongli Tan 외

Image matching, which aims to identify corresponding pixel locations between images, is crucial in a wide range of scientific disciplines, aiding in image registration, fusion, and analysis. In recent years, deep learnin…

Image Registration

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

2026-05-12 · Zhuoyu Cai, Dou Quan, Ning Huyan, Pei He 외 arxiv

Existing deep learning-based methods can capture shared features from optical and synthetic aperture radar (SAR) images for spatial alignment. However, optical-SAR registration remains challenging under large geometric d…

Image Registration

CM-Bench: A Comprehensive Cross-Modal Feature Matching Benchmark Bridging Visible and Infrared Images

2026-03-13 · Liangzheng Sun, Mengfan He, Xingyu Shao, Binbin Li 외 arxiv

Infrared-visible (IR-VIS) feature matching plays an essential role in cross-modality visual localization, navigation and perception. Along with the rapid development of deep learning techniques, a number of representativ…

Homography EstimationVisual LocalizationPose EstimationImage Matching

Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization

2025-09-16 · Yujia Lin, Nicholas Evans arxiv

Ensuring accurate localization of robots in environments without GPS capability is a challenging task. Visual Place Recognition (VPR) techniques can potentially achieve this goal, but existing RGB-based methods are sensi…

Visual Place RecognitionContrastive LearningGeometric Matching