paper-with-me

홈 › Papers

TerraMind: Large-Scale Generative Multimodality for Earth Observation

2025-04-15 · Johannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer, Rocco Sedona, Stefano Maurogiovanni, Jente Bosmans, Nikolaos Dionelis, Valerio Marsocci, Niklas Kopp, Rahul Ramachandran, Paolo Fraccaro, Thomas Brunschwiler, Gabriele Cavallaro, Juan Bernabe-Moreno, Nicolas Longépé

We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO). Unlike other multimodal models, TerraMind is pretrained on dual-scale representations combining both token-level and pixel-level data across modalities. On a token level, TerraMind encodes high-level contextual information to learn cross-modal relationships, while on a pixel level, TerraMind leverages fine-grained representations to capture critical spatial nuances. We pretrained TerraMind on nine geospatial modalities of a global, large-scale dataset. In this paper, we demonstrate that (i) TerraMind's dual-scale early fusion approach unlocks a range of zero-shot and few-shot applications for Earth observation, (ii) TerraMind introduces "Thinking-in-Modalities" (TiM) -- the capability of generating additional artificial data during finetuning and inference to improve the model output -- and (iii) TerraMind achieves beyond state-of-the-art performance in community-standard benchmarks for EO like PANGAEA. The pretraining dataset, the model weights, and our code is open-sourced under a permissive license.

📄 PDF Abstract BibTeX arXiv:2504.11171

Code (0)

등록된 구현이 없습니다.

Tasks

Earth Observation

Similar Papers 제목 키워드 기반

Ecological mapping with geospatial foundation models

2026-02-11 · Craig Mahlasi, Gciniwe S. Baloyi, Zaheed Gaffoor, Levente Klein 외 arxiv

The value of Earth observation foundation models for high-impact ecological applications remains insufficiently characterized. This study is one of the first to systematically evaluate the performance, limitations and pr…

EO-VAE: Towards A Multi-sensor Tokenizer for Earth Observation Data

2026-02-12 · Nils Lehmann, Yi Wang, Zhitong Xiong, Xiaoxiang Zhu arxiv

State-of-the-art generative image and video models rely heavily on tokenizers that compress high-dimensional inputs into more efficient latent representations. While this paradigm has revolutionized RGB generation, Earth…

Leveraging AI multimodal geospatial foundation models for improved near-real-time flood mapping at a global scale

2025-11-27 · Mirela G. Tulbure, Julio Caineta, Mark Broich, Mollie D. Gaines 외 arxiv

Floods are among the most damaging weather-related hazards, and in 2024, the warmest year on record, extreme flood events affected communities across five continents. Earth observation (EO) satellites provide critical, f…

MetaEarth3D: Unlocking World-scale 3D Generation with Spatially Scalable Generative Modeling

2026-04-19 · Jinqi Cao, Zhiping Yu, Baihong Lin, Chenyang Liu 외 arxiv

Recent generative AI models have achieved remarkable breakthroughs in language and visual understanding. However, although these models can generate realistic visual content, their spatial scale remains confined to bound…

3D Generation

Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction

2026-05-09 · Jinxiao Zhang, Runmin Dong, Xiyong Wu, Xihan Huang 외 arxiv

Earth observation is becoming one of the largest data-producing activities in science, yet current pipelines still treat compression as a storage and transmission tool rather than a new way to use data. We present a gene…