SatSwinMAE: Efficient Autoencoding for Multiscale Time-series Satellite Imagery
Recent advancements in foundation models have significantly impacted various fields, including natural language processing, computer vision, and multi-modal tasks. One area that stands to benefit greatly is Earth observation, where these models can efficiently process large-scale, unlabeled geospatial data. In this work we extend the SwinMAE model to integrate temporal information for satellite time-series data. The architecture employs a hierarchical 3D Masked Autoencoder (MAE) with Video Swin Transformer blocks to effectively capture multi-scale spatio-temporal dependencies in satellite imagery. To enhance transfer learning, we incorporate both encoder and decoder pretrained weights, along with skip connections to preserve scale-specific information. This forms an architecture similar to SwinUNet with an additional temporal component. Our approach shows significant performance improvements over existing state-of-the-art foundation models for all the evaluated downstream tasks: land cover segmentation, building density prediction, flood mapping, wildfire scar mapping and multi-temporal crop segmentation. Particularly, in the land cover segmentation task of the PhilEO Bench dataset, it outperforms other geospatial foundation models with a 10.4% higher accuracy.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderEarth ObservationRepresentation LearningSegmentationTime SeriesTransfer LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
HashEncoding: Autoencoding with Multiscale Coordinate Hashing
We present HashEncoding, a novel autoencoding architecture that leverages a non-parametric multiscale coordinate hash function to facilitate a per-pixel decoder without convolutions. By leveraging the space-folding behav…
DecoderOptical Flow EstimationTiMo: Spatiotemporal Foundation Model for Satellite Image Time Series
Satellite image time series (SITS) provide continuous observations of the Earth's surface, making them essential for applications such as environmental management and disaster assessment. However, existing spatiotemporal…
Temporal SequencesTime SeriesMSTCGAN: Multiscale time conditional generative adversarial network for long-term satellite image sequence prediction
Satellite image sequence prediction is a crucial and challenging task. Previous studies leverage optical flow methods or existing deep learning methods on spatial–temporal sequence models for the task. However, they suff…
Generative Adversarial NetworkOptical Flow EstimationPredictionLLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting
Time series forecasting remains a challenging task, particularly in the context of complex multiscale temporal patterns. This study presents LLM-Mixer, a framework that improves forecasting accuracy through the combinati…
Time SeriesTime Series ForecastingGIST: Towards Photorealistic Style Transfer via Multiscale Geometric Representations
State-of-the-art Style Transfer methods often leverage pre-trained encoders optimized for discriminative tasks, which may not be ideal for image synthesis. This can result in significant artifacts and loss of photorealis…
Image GenerationStyle Transfer