paper-with-me

홈 › Papers

Global–Local Information Soft-Alignment for Cross-Modal Remote-Sensing Image–Text Retrieval

2024-05-14 · journal 2024 5 · Gang Hu, Zaidao Wen, Yafei Lv, Jianting Zhang, Qian Wu

Cross-modal remote-sensing image–text retrieval (CMRSITR) is a challenging task that aims to retrieve target remote-sensing (RS) images based on textual descriptions. However, the modal gap between texts and RS images poses a significant challenge. RS images comprise multiple targets and complex backgrounds, necessitating the mining of both global and local information (GaLR) for effective CMRSITR. Existing approaches primarily focus on local image features while disregarding the local features of the text and their correspondence. These methods typically fuse global and local image features and align them with global text features. However, they struggle to eliminate the influence of cluttered backgrounds and may overlook crucial targets. To address these limitations, we propose a novel framework for CMRSITR based on a transformer architecture, which leverages global–local information soft alignment (GLISA) to enhance retrieval performance. Our framework incorporates a global image extraction module, which captures the global semantic features of image–text pairs and effectively represents the relationships among multiple targets in RS images. In addition, we introduce an adaptive local information extraction (ALIE) module that adaptively mines discriminative local clues from both RS images and texts, aligning the corresponding fine-grained information. To mitigate semantic ambiguities during the alignment of local features, we design a local information soft-alignment (LISA) module. In comparative evaluations using two public CMRSITR datasets, our proposed method achieves state-of-the-art results, surpassing not only traditional cross-modal retrieval methods by a substantial margin but also other contrastive language-image pretraining (CLIP)-based methods.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalCross-Modal Retrieval on RSITMDImage-text RetrievalRetrievalText Retrieval

Methods 이 논문이 사용한 방법론

Focus 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

RNAmountAlign: efficient software for local, global, semiglobal pairwise and multiple RNA sequence/structure alignment

2018-08-10

Alignment of structural RNAs is an important problem with a wide range of applications. Since function is often determined by molecular structure, RNA alignment programs should take into account both sequence and base-pa…

Benchmarking

3D PersonVLAD: Learning Deep Global Representations for Video-based Person Re-identification

2018-12-26 · Lin Wu, Yang Wang, Ling Shao, Meng Wang

In this paper, we introduce a global video representation to video-based person re-identification (re-ID) that aggregates local 3D features across the entire video extent. Most of the existing methods rely on 2D convolut…

Person Re-IdentificationVideo-Based Person Re-Identification

TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local Correspondence

2025-06-26 · Feng Jiang, Mangal Prakash, Hehuan Ma, Jianyuan Deng 외

Molecular property prediction aims to learn representations that map chemical structures to functional properties. While multimodal learning has emerged as a powerful paradigm to learn molecular representations, prior wo…

Molecular Property Predictionmolecular representationProperty PredictionRepresentation Learning

Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval

2024-11-22 · Zengbao Sun, Ming Zhao, Gaorui Liu, André Kaup

Remote sensing cross-modal text-image retrieval (RSCTIR) has gained attention for its utility in information mining. However, challenges remain in effectively integrating global and local information due to variations in…

Image RetrievalRerankingRetrievalText Retrieval+1

AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition

2024-10-21 · Zehua Liu, Xiaolou Li, Chen Chen, Li Guo 외

Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip movements, VSR tasks require effectively ut…

cross-modal alignmentspeech-recognitionSpeech RecognitionVisual Speech Recognition