paper-with-me

홈 › Papers

SatSwinMAE: Efficient Autoencoding for Multiscale Time-series Satellite Imagery

2024-05-03 · Yohei Nakayama, Jiawei Su, Luis M. Pazos-Outón

Recent advancements in foundation models have significantly impacted various fields, including natural language processing, computer vision, and multi-modal tasks. One area that stands to benefit greatly is Earth observation, where these models can efficiently process large-scale, unlabeled geospatial data. In this work we extend the SwinMAE model to integrate temporal information for satellite time-series data. The architecture employs a hierarchical 3D Masked Autoencoder (MAE) with Video Swin Transformer blocks to effectively capture multi-scale spatio-temporal dependencies in satellite imagery. To enhance transfer learning, we incorporate both encoder and decoder pretrained weights, along with skip connections to preserve scale-specific information. This forms an architecture similar to SwinUNet with an additional temporal component. Our approach shows significant performance improvements over existing state-of-the-art foundation models for all the evaluated downstream tasks: land cover segmentation, building density prediction, flood mapping, wildfire scar mapping and multi-temporal crop segmentation. Particularly, in the land cover segmentation task of the PhilEO Bench dataset, it outperforms other geospatial foundation models with a 10.4% higher accuracy.

📄 PDF Abstract BibTeX arXiv:2405.02512

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderEarth ObservationRepresentation LearningSegmentationTime SeriesTransfer Learning

Methods 이 논문이 사용한 방법론

Attention 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

HashEncoding: Autoencoding with Multiscale Coordinate Hashing

2022-11-29 · Lukas Zhornyak, Zhengjie Xu, Haoran Tang, Jianbo Shi

We present HashEncoding, a novel autoencoding architecture that leverages a non-parametric multiscale coordinate hash function to facilitate a per-pixel decoder without convolutions. By leveraging the space-folding behav…

DecoderOptical Flow Estimation

TiMo: Spatiotemporal Foundation Model for Satellite Image Time Series

2025-05-13 · Xiaolei Qin, Di Wang, Jing Zhang, Fengxiang Wang 외

Satellite image time series (SITS) provide continuous observations of the Earth's surface, making them essential for applications such as environmental management and disaster assessment. However, existing spatiotemporal…

Temporal SequencesTime Series

MSTCGAN: Multiscale time conditional generative adversarial network for long-term satellite image sequence prediction

2022-06-01 · IEEE Transactions on Geoscience and Remote Sensing 2022 6 · Kuai Dai, Xutao Li, Yunming Ye, Shanshan Feng 외

Satellite image sequence prediction is a crucial and challenging task. Previous studies leverage optical flow methods or existing deep learning methods on spatial–temporal sequence models for the task. However, they suff…

Generative Adversarial NetworkOptical Flow EstimationPrediction

LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting

2024-10-15 · Md Kowsher, Md. Shohanur Islam Sobuj, Nusrat Jahan Prottasha, E. Alejandro Alanis 외

Time series forecasting remains a challenging task, particularly in the context of complex multiscale temporal patterns. This study presents LLM-Mixer, a framework that improves forecasting accuracy through the combinati…

Time SeriesTime Series Forecasting

GIST: Towards Photorealistic Style Transfer via Multiscale Geometric Representations

2024-12-03 · Renan A. Rojas-Gomez, Minh N. Do

State-of-the-art Style Transfer methods often leverage pre-trained encoders optimized for discriminative tasks, which may not be ideal for image synthesis. This can result in significant artifacts and loss of photorealis…

Image GenerationStyle Transfer