paper-with-me

홈 › Papers

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

2026-05-12 · Zhuoyu Cai, Dou Quan, Ning Huyan, Pei He, Shuang Wang, Licheng Jiao arxiv

Existing deep learning-based methods can capture shared features from optical and synthetic aperture radar (SAR) images for spatial alignment. However, optical-SAR registration remains challenging under large geometric deformations, because the model needs to simultaneously handle cross-modal appearance discrepancies and complex spatial transformations. To address this issue, this paper proposes a text semantic-assisted cross-modal image registration framework, named TAR, for optical and SAR images. TAR exploits text semantic priors from remote sensing scenes and land-cover categories to alleviate the modality gap and enhance cross-modal feature learning. TAR consists of three components: a multi-scale visual feature learning (MSFL) module, a text-assisted feature enhancement (TAFE) module, and a coarse-to-fine dense matching (CFDM) module. MSFL extracts multi-scale visual features from optical and SAR images. TAFE constructs text descriptors related to remote sensing scenes and land-cover objects, and uses a frozen RemoteCLIP text encoder to extract text features. These text features are introduced through visual-text interaction to enhance high-level visual features for more reliable coarse matching. CFDM then establishes coarse correspondences based on the enhanced high-level features and refines the matched locations using low-level features. Experimental results on cross-modal remote sensing images demonstrate the effectiveness of TAR, which achieves stronger matching performance than several state-of-the-art methods and yields significant gains under large geometric deformations.

📄 PDF Abstract BibTeX arXiv:2605.12064

Code (0)

등록된 구현이 없습니다.

Tasks

Image Registration

Similar Papers 제목 키워드 기반

Language-Assisted Image Clustering Guided by Discriminative Relational Signals and Adaptive Semantic Centers

2026-03-25 · Jun Ma, Xu Zhang, Zhengxing Jiao, Yaxin Hou 외 arxiv

Language-Assisted Image Clustering (LAIC) augments the input images with additional texts with the help of vision-language models (VLMs) to promote clustering performance. Despite recent progress, existing LAIC methods o…

Image Clustering

OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment

2025-10-15 · Rongjun Chen, Chengsi Yao, Jinchang Ren, Xianxian Zeng 외 arxiv

Text-image alignment constitutes a foundational challenge in multimedia content understanding, where effective modeling of cross-modal semantic correspondences critically enhances retrieval system performance through joi…

Cross-Modal Retrieval

GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates

2026-01-31 · Xingyu Luo, Yidong Cai, Jie Liu, Jie Tang 외 arxiv

Vision-language tracking has gained increasing attention in many scenarios. This task simultaneously deals with visual and linguistic information to localize objects in videos. Despite its growing utility, the developmen…

Visual Tracking

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding

2026-05-05 · Ran Ran, Jiwei Wei, Shuchang Zhou, Yitong Qin 외 arxiv

Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while directly matching the query to moments results in insufficient discr…

A scoping review on multimodal deep learning in biomedical images and texts

2023-07-14 · Zhaoyi Sun, Mingquan Lin, Qingqing Zhu, Qianqian Xie 외

Computer-assisted diagnostic and prognostic systems of the future should be capable of simultaneously processing multimodal data. Multimodal deep learning (MDL), which involves the integration of multiple sources of data…

Cross-Modal RetrievalDecision MakingDiagnosticMultimodal Deep Learning+4