paper-with-me

Papers

GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding

2025-12-02 · Jiaqi Liu, Ronghao Fu, Haoran Liu, Lang Sun, Bo Yang arxiv

Autoregressive models are structurally misaligned with the inherently parallel nature of geospatial understanding, forcing a rigid sequential narrative onto scenes and fundamentally hindering the generation of structured and coherent outputs. We challenge this paradigm by reframing geospatial generation as a parallel refinement process, enabling a holistic, coarse-to-fine synthesis that resolves all semantic elements simultaneously. To operationalize this, we introduce GeoDiT, the first diffusion-based vision-language model tailored for the geospatial domain. Extensive experiments demonstrate that GeoDiT establishes a new state-of-the-art on benchmarks requiring structured, object-centric outputs. It achieves significant gains in image captioning, visual grounding, and multi-object detection, precisely the tasks where autoregressive models falter. Our work validates that aligning the generative process with the data's intrinsic structure is key to unlocking superior performance in complex geospatial analysis.

📄 PDF Abstract BibTeX arXiv:2512.02505

Code (0)

등록된 구현이 없습니다.

Tasks

Visual GroundingObject DetectionImage Captioning

Similar Papers 제목 키워드 기반

Textual Supervision Enhances Geospatial Representations in Vision-Language Models

2026-06-05 · Marcelo Sartori Locatelli, Fernando Tonucci, Jea Kwon, Luiz Felipe Vecchietti 외 arxiv

Geospatial understanding is a critical yet underexplored dimension in the development of machine learning systems for tasks such as image geolocation and spatial reasoning. In this work, we analyze the geospatial represe…

Spatial Reasoning

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

2024-11-28 · Muhammad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Kartik Kuckreja 외

While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they fall short in addressing the unique demands of geospatial applications. Generic VLM benchmarks are not designed to handle t…

BenchmarkingObject CountingScene Understanding

Observing Health Outcomes Using Remote Sensing Imagery and Geo-Context Guided Visual Transformer

2026-01-26 · Yu Li, Guilherme N. DeSouza, Praveen Rao, Chi-Ren Shyu arxiv

Visual transformers have driven major progress in remote sensing image analysis, particularly in object detection and segmentation. Recent vision-language and multimodal models further extend these capabilities by incorp…

Object Detection

SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts

2025-12-02 · Jiaqi Liu, Ronghao Fu, Lang Sun, Haoran Liu 외 arxiv

The emergence of large vision-language models (VLMs) has significantly enhanced the efficiency and flexibility of geospatial interpretation. However, general-purpose VLMs remain suboptimal for remote sensing (RS) tasks. …

Representation Learning

GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes

2025-11-27 · Di Wang, Shunyu Liu, Wentao Jiang, Fengxiang Wang 외 arxiv

Multimodal large language models (MLLMs) have undergone rapid development in advancing geospatial scene understanding. Recent studies have sought to enhance the reasoning capabilities of remote sensing MLLMs, typically t…

Reinforcement LearningScene Understanding