paper-with-me

홈 › Papers

MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

2026-01-13 · Aditya Chaudhary, Sneha Barman, Mainak Singha, Ankit Jha, Girish Mishra, Biplab Banerjee arxiv

In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language semantics using vision-language models such as CLIP. With the increasing availability of multimodal Earth observation data, there is a growing need for methods that effectively fuse spectral, spatial, and geometric information while enabling semantic-level understanding. MMLGNet employs modality-specific encoders and aligns visual features with handcrafted textual embeddings in a shared latent space via bi-directional contrastive learning. Inspired by CLIP's training paradigm, our approach bridges the gap between high-dimensional remote sensing data and language-guided interpretation. Notably, MMLGNet achieves strong performance with simple CNN-based encoders, outperforming several established multimodal visual-only methods on two benchmark datasets, demonstrating the significant benefit of language supervision. Codes are available at https://github.com/AdityaChaudhary2913/CLIP_HSI.

📄 PDF Abstract BibTeX arXiv:2601.08420

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image Segmentation

2024-09-20 · Sen Lei, Xinyu Xiao, Tianlin Zhang, Heng-Chao Li 외

Given a language expression, referring remote sensing image segmentation (RRSIS) aims to identify ground objects and assign pixel-wise labels within the imagery. The one of key challenges for this task is to capture disc…

Image SegmentationReferring ExpressionSemantic Segmentation

Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval

2024-11-22 · Zengbao Sun, Ming Zhao, Gaorui Liu, André Kaup

Remote sensing cross-modal text-image retrieval (RSCTIR) has gained attention for its utility in information mining. However, challenges remain in effectively integrating global and local information due to variations in…

Image RetrievalRerankingRetrievalText Retrieval+1

Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval

2024-05-29 · Rui Yang, Shuang Wang, Yingping Han, Yuanheng Li 외

Remote Sensing Image-Text Retrieval (RSITR) is pivotal for knowledge services and data mining in the remote sensing (RS) domain. Considering the multi-scale representations in image content and text vocabulary can enable…

cross-modal alignmentImage-text RetrievalRetrievalText Retrieval

GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning

2026-03-10 · Xiao Yang, Ronghao Fu, Zhuoran Duan, Zhiwen Lin 외 arxiv

Vision-language pretraining models have made significant progress in bridging remote sensing imagery with natural language. However, existing approaches often fail to effectively integrate multi-granular visual and textu…

Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism

2025-12-29 · Siyu Zhang, Lianlei Shan, Runhe Qiu arxiv

Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant applica…

Computational EfficiencyInformation ExtractionCross-Modal RetrievalImage Captioning