paper-with-me

홈 › Papers

Probing Geospatial SSL Representations with Environmental Signals

2026-07-06 · Rohita Mocharla, Vishal M. Patel arxiv

Self-supervised learning (SSL) is designed to learn generic, transferable representations rather than representations optimized for a single task. Most geospatial benchmarks evaluate representations solely through downstream tasks, providing limited insight into the information encoded within the representation itself. We ask a different question: do SSL representations of satellite imagery preserve statistical associations with environmental variables that co-vary with the imaging process? To answer this question, we probe SSL representations using co-located ERA5 reanalysis variables, a global dataset of physically consistent environmental variables, including temperature, precipitation, surface solar radiation, surface pressure, and volumetric soil water. These variables are physically related to the spectral reflectance and radar backscatter recorded by Sentinel-1 and Sentinel-2, making them meaningful evaluation targets despite not being used during SSL pretraining. We complement this probing analysis with intrinsic representation metrics to characterize representation geometry and investigate how these properties relate to downstream performance and the encoding of environmental signals. Using DINO, MAE, and MoCo models trained under identical conditions, we show that representation-level metrics distinguish models with similar downstream benchmark performance, providing complementary information beyond task-driven benchmarks. We further find that the linear accessibility of environmental signals is associated with performance on environmentally dependent tasks in the PANGAEA benchmark. Finally, we release ERA5 annotations co-located with the SSL4EO dataset to enable physically grounded representation evaluation for future geospatial foundation models.

📄 PDF Abstract BibTeX arXiv:2607.05207

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation

2026-05-14 · Yuhao Liu, Sadeer Al-Kindi, Ashok Veeraraghavan, Guha Balakrishnan arxiv

Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing geospatial foundation models do not directly model the structured soc…

More than Correlation: Do Large Language Models Learn Causal Representations of Space?

2023-12-26 · Yida Chen, Yixian Gan, Sijia Li, Li Yao 외

Recent work found high mutual information between the learned representations of large language models (LLMs) and the geospatial property of its input, hinting an emergent internal model of space. However, whether this i…

Beyond AlphaEarth: Toward Human-Centered Geospatial Foundation Models via POI-Guided Contrastive Learning

2025-10-10 · Junyuan Liu, Quan Qin, Guangsheng Dong, Xinglei Wang 외 arxiv

Recent geospatial foundation models (GFMs) produce spatially extensive representations of the Earth's surface that capture rich physical and environmental patterns. Among them, the AlphaEarth Foundation (AE) represents a…

Natural Language QueriesRepresentation LearningContrastive Learning

Spatial Representation Learning Beyond Pixels: Unifying Raster Data and Vector Semantics for Human-Centric Geospatial Foundation Models

2026-06-01 · Steffen Knoblauch, Hao Li, Gengchen Mai, Konstantin Klemmer 외 arxiv

Earth Observation (EO) has fundamentally transformed the monitoring of environmental processes and human activities up to planetary scale. Recent advances in self-supervised learning have given rise to Earth Observation …

Self-Supervised LearningRepresentation Learning

Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context

2026-08-21 · Utsav Poudel, Jagannath Aryal, Subramaniyaswamy Vairavasundaram arxiv

Environmental exposures such as air pollution and greenness have been associated with affective and cognitive outcomes, but EEG and environmental datasets are rarely jointly georeferenced. We investigate whether literatu…