paper-with-me

홈 › Papers

RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation

2025-12-04 · Nicolas Houdré, Diego Marcos, Hugo Riffaud de Turckheim, Dino Ienco, Laurent Wendling, Camille Kurtz, Sylvain Lobry arxiv

Earth observation (EO) data spans a wide range of spatial, spectral, and temporal resolutions, from high-resolution optical imagery to low resolution multispectral products or radar time series. While recent foundation models have improved multimodal integration for learning meaningful representations, they often expect fixed input resolutions or are based on sensor-specific encoders limiting generalization across heterogeneous EO modalities. To overcome these limitations we introduce RAMEN, a resolution-adjustable multimodal encoder that learns a shared visual representation across EO data in a fully sensor-agnostic manner. RAMEN treats the modality and spatial and temporal resolutions as key input data features, enabling coherent analysis across modalities within a unified latent space. Its main methodological contribution is to define spatial resolution as a controllable output parameter, giving users direct control over the desired level of detail at inference and allowing explicit trade-offs between spatial precision and computational cost. We train a single, unified transformer encoder reconstructing masked multimodal EO data drawn from diverse sources, ensuring generalization across sensors and resolutions. Once pretrained, RAMEN transfers effectively to both known and unseen sensor configurations and outperforms larger state-of-the-art models on the community-standard PANGAEA benchmark, containing various multi-sensor and multi-resolution downstream tasks. Our code and pretrained model are available at https://github.com/nicolashoudre/RAMEN.

📄 PDF Abstract BibTeX arXiv:2512.05025

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation

2026-06-22 · Yohann Perron, Guillaume Astruc, Nicolas Gonthier, Clement Mallet 외 arxiv

Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfer to Earth Observation (EO), where input modalities, scales, and resolutions vary widely. We introduce …

AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities

2024-12-18 · CVPR 2025 1 · Guillaume Astruc, Nicolas Gonthier, Clement Mallet, Loic Landrieu

Geospatial models must adapt to the diversity of Earth observation data in terms of resolutions, scales, and modalities. However, existing approaches expect fixed input configurations, which limits their practical applic…

Change DetectionDiversityEarth Observation

Charon: a FrameNet Annotation Tool for Multimodal Corpora

2022-05-24 · LREC (LAW) 2022 6 · Frederico Belcavello, Marcelo Viridiano, Ely Edison Matos, Tiago Timponi Torrent

This paper presents Charon, a web tool for annotating multimodal corpora with FrameNet categories. Annotation can be made for corpora containing both static images and video sequences paired - or not - with text sequence…

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

2026-03-07 · Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati 외 arxiv

We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scali…

EarthScape: A Multimodal Dataset for Surficial Geologic Mapping and Earth Surface Analysis

2025-03-19 · Matthew Massey, Abdullah-Al-Zubaer Imran

Surficial geologic mapping is essential for understanding Earth surface processes, addressing modern challenges such as climate change and national security, and supporting common applications in engineering and resource…