paper-with-me

Papers

Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?

2025-03-10 · Yuru Jia, Valerio Marsocci, Ziyang Gong, Xue Yang, Maarten Vergauwen, Andrea Nascetti

Self-supervised learning (SSL) has revolutionized representation learning in Remote Sensing (RS), advancing Geospatial Foundation Models (GFMs) to leverage vast unlabeled satellite imagery for diverse downstream tasks. Currently, GFMs primarily focus on discriminative objectives, such as contrastive learning or masked image modeling, owing to their proven success in learning transferable representations. However, generative diffusion models--which demonstrate the potential to capture multi-grained semantics essential for RS tasks during image generation--remain underexplored for discriminative applications. This prompts the question: can generative diffusion models also excel and serve as GFMs with sufficient discriminative power? In this work, we answer this question with SatDiFuser, a framework that transforms a diffusion-based generative geospatial foundation model into a powerful pretraining tool for discriminative RS. By systematically analyzing multi-stage, noise-dependent diffusion features, we develop three fusion strategies to effectively leverage these diverse representations. Extensive experiments on remote sensing benchmarks show that SatDiFuser outperforms state-of-the-art GFMs, achieving gains of up to +5.7% mIoU in semantic segmentation and +7.9% F1-score in classification, demonstrating the capacity of diffusion-based generative foundation models to rival or exceed discriminative GFMs. Code will be released.

📄 PDF Abstract BibTeX arXiv:2503.07890

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningImage GenerationRepresentation LearningSelf-Supervised LearningSemantic Segmentation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Contrastive Learning 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding

2025-12-02 · Jiaqi Liu, Ronghao Fu, Haoran Liu, Lang Sun 외 arxiv

Autoregressive models are structurally misaligned with the inherently parallel nature of geospatial understanding, forcing a rigid sequential narrative onto scenes and fundamentally hindering the generation of structured…

Visual GroundingObject DetectionImage Captioning

GeoEvolve: Automating Geospatial Model Discovery via Multi-Agent Large Language Models

2025-09-25 · Peng Luo, Xiayin Lou, Yu Zheng, Zhuo Zheng 외 arxiv

Geospatial modeling provides critical solutions for pressing global challenges such as sustainability and climate change. Existing large language model (LLM)-based algorithm discovery frameworks, such as AlphaEvolve, exc…

Generative Adversarial Models for Extreme Geospatial Downscaling

2024-02-21 · Guiye Li, Guofeng Cao

Addressing the challenges of climate change requires accurate and high-resolution mapping of geospatial data, especially climate and weather variables. However, many existing geospatial datasets, such as the gridded outp…

Image Super-ResolutionSuper-Resolution

GeoResponder: Towards Building Geospatial LLMs for Time-Critical Disaster Response

2025-09-18 · Ahmed El Fekih Zguir, Ferda Ofli, Muhammad Imran arxiv

LLMs excel at linguistic tasks but lack the inner geospatial capabilities needed for time-critical disaster response, where reasoning about road networks, coordinates, and access to essential infrastructure such as hospi…

Spatial Reasoning

MapQA: Open-domain Geospatial Question Answering on Map Data

2025-03-10 · Zekun Li, Malcolm Grossman, Eric, Qasemi 외

Geospatial question answering (QA) is a fundamental task in navigation and point of interest (POI) searches. While existing geospatial QA datasets exist, they are limited in both scale and diversity, often relying solely…

DiversityLanguage ModelingLanguage ModellingLarge Language Model+2