Survey of Multimodal Geospatial Foundation Models: Techniques, Applications, and Challenges
Foundation models have transformed natural language processing and computer vision, and their impact is now reshaping remote sensing image analysis. With powerful generalization and transfer learning capabilities, they align naturally with the multimodal, multi-resolution, and multi-temporal characteristics of remote sensing data. To address unique challenges in the field, multimodal geospatial foundation models (GFMs) have emerged as a dedicated research frontier. This survey delivers a comprehensive review of multimodal GFMs from a modality-driven perspective, covering five core visual and vision-language modalities. We examine how differences in imaging physics and data representation shape interaction design, and we analyze key techniques for alignment, integration, and knowledge transfer to tackle modality heterogeneity, distribution shifts, and semantic gaps. Advances in training paradigms, architectures, and task-specific adaptation strategies are systematically assessed alongside a wealth of emerging benchmarks. Representative multimodal visual and vision-language GFMs are evaluated across ten downstream tasks, with insights into their architectures, performance, and application scenarios. Real-world case studies, spanning land cover mapping, agricultural monitoring, disaster response, climate studies, and geospatial intelligence, demonstrate the practical potential of GFMs. Finally, we outline pressing challenges in domain generalization, interpretability, efficiency, and privacy, and chart promising avenues for future research.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain GeneralizationTransfer LearningSimilar Papers 제목 키워드 기반
Towards Vision-Language Geo-Foundation Model: A Survey
Vision-Language Foundation Models (VLFMs) have made remarkable progress on various multimodal tasks, such as image captioning, image-text retrieval, visual question answering, and visual grounding. However, most methods …
Earth ObservationImage CaptioningImage-text Retrievalmodel+5Deep Learning Techniques for Geospatial Data Analysis
Consumer electronic devices such as mobile handsets, goods tagged with RFID labels, location and position sensors are continuously generating a vast amount of location enriched data called geospatial data. Conventionally…
Deep Learningimage-classificationImage ClassificationObject Recognition+1Self-supervised Learning for Geospatial AI: A Survey
The proliferation of geospatial data in urban and territorial environments has significantly facilitated the development of geospatial artificial intelligence (GeoAI) across various urban applications. Given the vast yet…
Self-Supervised LearningSurveyAgentic AI in Remote Sensing: Foundations, Taxonomy, and Emerging Systems
The paradigm of Earth Observation analysis is shifting from static deep learning models to autonomous agentic AI. Although recent vision foundation models and multimodal large language models advance representation learn…
Representation LearningOn the Opportunities and Challenges of Foundation Models for Geospatial Artificial Intelligence
Large pre-trained models, also known as foundation models (FMs), are trained in a task-agnostic manner on large-scale data and can be adapted to a wide range of downstream tasks by fine-tuning, few-shot, or even zero-sho…
Few-Shot LearningScene ClassificationTime Series ForecastingToponym Recognition+1