paper-with-me

Papers

Efficient Remote Sensing with Harmonized Transfer Learning and Modality Alignment

2024-04-28 · Tengjun Huang

With the rise of Visual and Language Pretraining (VLP), an increasing number of downstream tasks are adopting the paradigm of pretraining followed by fine-tuning. Although this paradigm has demonstrated potential in various multimodal downstream tasks, its implementation in the remote sensing domain encounters some obstacles. Specifically, the tendency for same-modality embeddings to cluster together impedes efficient transfer learning. To tackle this issue, we review the aim of multimodal transfer learning for downstream tasks from a unified perspective, and rethink the optimization process based on three distinct objectives. We propose "Harmonized Transfer Learning and Modality Alignment (HarMA)", a method that simultaneously satisfies task constraints, modality alignment, and single-modality uniform alignment, while minimizing training overhead through parameter-efficient fine-tuning. Remarkably, without the need for external data for training, HarMA achieves state-of-the-art performance in two popular multimodal retrieval tasks in the field of remote sensing. Our experiments reveal that HarMA achieves competitive and even superior performance to fully fine-tuned models with only minimal adjustable parameters. Due to its simplicity, HarMA can be integrated into almost all existing multimodal pretraining models. We hope this method can facilitate the efficient application of large models to a wide range of downstream tasks while significantly reducing the resource consumption. Code is available at https://github.com/seekerhuang/HarMA.

📄 PDF Abstract BibTeX arXiv:2404.18253

Code (1)

seekerhuang/harma 공식 구현 pytorch

Tasks

Cross-Modal RetrievalImage RetrievalImage-to-Text Retrievalparameter-efficient fine-tuningRetrievalText RetrievalTransfer Learning

Similar Papers 제목 키워드 기반

Contrastive Parameter Disentanglement for Multi-modal Remote Sensing Image Generation

2026-07-26 · Yu Zhang, Wenda Zhao, Haojun Tang, Haipeng Wang arxiv

Existing remote sensing image generation methods are largely confined to single-modality synthesis and therefore fail to exploit the complementary information inherent in multimodal imagery. To address this limitation, w…

Image Generation

GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing

2026-04-12 · Maram Hasan, Md Aminur Hossain, Savitra Roy, Souparna Bhowmik 외 arxiv

Effective foundation modeling in remote sensing requires spatially aligned heterogeneous modalities coupled with semantically grounded supervision, yet such resources remain limited at scale. We present GeoMeld, a large-…

Representation Learning

Rethinking Efficient Mixture-of-Experts for Remote Sensing Modality-Missing Classification

2025-11-14 · Qinghao Gao, Jiahui Qu, Wenqian Dong arxiv

Multimodal remote sensing classification often suffers from missing modalities caused by sensor failures and environmental interference, leading to severe performance degradation. In this work, we rethink missing-modalit…

A Survey on Remote Sensing Foundation Models: From Vision to Multimodality

2025-03-28 · Ziyue Huang, Hongxi Yan, Qiqi Zhan, Shuai Yang 외

The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data interpretation. These models combine variou…

Change DetectionLand Cover ClassificationTransfer Learning

STARS: Shared-specific Translation and Alignment for missing-modality Remote Sensing Semantic Segmentation

2026-01-24 · Tong Wang, Xiaodong Zhang, Guanzhou Chen, Jiaqi Wang 외 arxiv

Multimodal remote sensing technology significantly enhances the understanding of surface semantics by integrating heterogeneous data such as optical images, Synthetic Aperture Radar (SAR), and Digital Surface Models (DSM…

Semantic Segmentation