RS-DFM: A Remote Sensing Distributed Foundation Model for Diverse Downstream Tasks
Remote sensing lightweight foundation models have achieved notable success in online perception within remote sensing. However, their capabilities are restricted to performing online inference solely based on their own observations and models, thus lacking a comprehensive understanding of large-scale remote sensing scenarios. To overcome this limitation, we propose a Remote Sensing Distributed Foundation Model (RS-DFM) based on generalized information mapping and interaction. This model can realize online collaborative perception across multiple platforms and various downstream tasks by mapping observations into a unified space and implementing a task-agnostic information interaction strategy. Specifically, we leverage the ground-based geometric prior of remote sensing oblique observations to transform the feature mapping from absolute depth estimation to relative depth estimation, thereby enhancing the model's ability to extract generalized features across diverse heights and perspectives. Additionally, we present a dual-branch information compression module to decouple high-frequency and low-frequency feature information, achieving feature-level compression while preserving essential task-agnostic details. In support of our research, we create a multi-task simulation dataset named AirCo-MultiTasks for multi-UAV collaborative observation. We also conduct extensive experiments, including 3D object detection, instance segmentation, and trajectory prediction. The numerous results demonstrate that our RS-DFM achieves state-of-the-art performance across various downstream tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object DetectionDepth EstimationInstance Segmentationobject-detectionObject DetectionSemantic SegmentationTrajectory PredictionSimilar Papers 제목 키워드 기반
Parameter Efficient Self-Supervised Geospatial Domain Adaptation
As large-scale foundation models become publicly available for different domains efficiently adapting them to individual downstream applications and additional data modalities has turned into a central challenge. For…
Domain AdaptationLinear evaluationRemoteCLIP: A Vision Language Foundation Model for Remote Sensing
General-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation model…
ClassificationCross-Modal Retrievalimage-classificationImage Classification+8A Billion-scale Foundation Model for Remote Sensing Images
As the potential of foundation models in visual tasks has garnered significant attention, pretraining these models before downstream tasks has become a crucial step. The three key factors in pretraining foundation models…
object-detectionObject DetectionObject Detection In Aerial ImagesSemantic SegmentationOne for All: Toward Unified Foundation Models for Earth Vision
Foundation models characterized by extensive parameters and trained on large-scale datasets have demonstrated remarkable efficacy across various downstream tasks for remote sensing data. Current remote sensing foundation…
AllTESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis
Satellite remote sensing (RS) enables a wide array of downstream Earth observation (EO) applications, including climate modeling, carbon accounting, and strategies for conservation and sustainable land use. We present TE…
Earth ObservationSelf-Supervised Learning