paper-with-me

홈 › Papers

Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval

2024-03-15 · Qian Wang, Jia-Chen Gu, Zhen-Hua Ling

Audio-text retrieval (ATR), which retrieves a relevant caption given an audio clip (A2T) and vice versa (T2A), has recently attracted much research attention. Existing methods typically aggregate information from each modality into a single vector for matching, but this sacrifices local details and can hardly capture intricate relationships within and between modalities. Furthermore, current ATR datasets lack comprehensive alignment information, and simple binary contrastive learning labels overlook the measurement of fine-grained semantic differences between samples. To counter these challenges, we present a novel ATR framework that comprehensively captures the matching relationships of multimodal information from different perspectives and finer granularities. Specifically, a fine-grained alignment method is introduced, achieving a more detail-oriented matching through a multiscale process from local to global levels to capture meticulous cross-modal relationships. In addition, we pioneer the application of cross-modal similarity consistency, leveraging intra-modal similarity relationships as soft supervision to boost more intricate alignment. Extensive experiments validate the effectiveness of our approach, outperforming previous methods by significant margins of at least 3.9% (T2A) / 6.9% (A2T) R@1 on the AudioCaps dataset and 2.9% (T2A) / 5.4% (A2T) R@1 on the Clotho dataset.

📄 PDF Abstract BibTeX arXiv:2403.10146

Code (0)

등록된 구현이 없습니다.

Tasks

AudioCapsContrastive LearningRetrievalText Retrieval

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Semi-supervised Multiscale Matching for SAR-Optical Image

2025-08-11 · Jingze Gai, Changchun Li arxiv

Driven by the complementary nature of optical and synthetic aperture radar (SAR) images, SAR-optical image matching has garnered significant interest. Most existing SAR-optical image matching methods aim to capture effec…

Image Matching

Balancing multiscale similarity and cartographic constraints: A similarity-driven optimization framework for line generalization

2026-07-28 · Pengbo Li, Haowen Yan, Xiaomin Lu, Binbin Lin arxiv

Cartographic generalization is essential for generating multiscale map representations by balancing information preservation and cartographic readability. However, automated generalization remains challenging because exi…

Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image Retrieval

2022-04-21 · Zhiqiang Yuan, Wenkai Zhang, Kun fu, Xuan Li 외

Remote sensing (RS) cross-modal text-image retrieval has attracted extensive attention for its advantages of flexible input and efficient query. However, traditional methods ignore the characteristics of multi-scale and …

Cross-Modal RetrievalImage RetrievalRetrievalSentence+1

Attention-Based Multimodal Image Matching

2021-03-20 · Aviad Moreshet, Yosi Keller

We propose an attention-based approach for multimodal image patch matching using a Transformer encoder attending to the feature maps of a multiscale Siamese CNN. Our encoder is shown to efficiently aggregate multiscale i…

Multimodal Patch MatchingPatch Matching

Cross-Domain Visual Matching via Generalized Similarity Measure and Feature Learning

2016-05-13 · Liang Lin, Guangrun Wang, WangMeng Zuo, Xiangchu Feng 외

Cross-domain visual data matching is one of the fundamental problems in many real-world vision tasks, e.g., matching persons across ID photos and surveillance videos. Conventional approaches to this problem usually invol…

Face VerificationModel OptimizationPerson Re-IdentificationRepresentation Learning