paper-with-me

홈 › Papers

RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning

2026-01-29 · Shiqi Huang, Shuting He, Bihan Wen arxiv

Remote Sensing Visual Grounding (RSVG) aims to localize target objects in large-scale aerial imagery based on natural language descriptions. Owing to the vast spatial scale and high semantic ambiguity of remote sensing scenes, these descriptions often rely heavily on positional cues, posing unique challenges for Multimodal Large Language Models (MLLMs) in spatial reasoning. To leverage this unique feature, we propose a reasoning-guided, position-aware post-training framework, dubbed \textbf{RSGround-R1}, to progressively enhance spatial understanding. Specifically, we first introduce Chain-of-Thought Supervised Fine-Tuning (CoT-SFT) using synthetically generated RSVG reasoning data to establish explicit position awareness. Reinforcement Fine-Tuning (RFT) is then applied, augmented by our newly designed positional reward that provides continuous and distance-aware guidance toward accurate localization. Moreover, to mitigate incoherent localization behaviors across rollouts, we introduce a spatial consistency guided optimization scheme that dynamically adjusts policy updates based on their spatial coherence, ensuring stable and robust convergence. Extensive experiments on RSVG benchmarks demonstrate superior performance and generalization of our model.

📄 PDF Abstract BibTeX arXiv:2601.21634

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningVisual Grounding

Similar Papers 제목 키워드 기반

Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting

2026-05-30 · Panav Shah, Geet Sethi, Ashutosh Gandhe arxiv

Visual grounding aims to locate image regions that correspond to natural language descriptions and is a key component of interpretable vision systems. In remote sensing imagery, grounding is particularly challenging due …

Visual Grounding

Think and Answer ME: Benchmarking and Exploring Multi-Entity Reasoning Grounding in Remote Sensing

2026-03-13 · Shuchang Lyu, Haiquan Wen, Guangliang Cheng, Meng Li 외 arxiv

Recent advances in reasoning language models and reinforcement learning with verifiable rewards have significantly enhanced multi-step reasoning capabilities. This progress motivates the extension of reasoning paradigms …

Reinforcement LearningVisual Grounding

GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

2025-12-02 · Peirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin 외 arxiv

Recent advances in multimodal large language models(MLLMs) have led to remarkable progress in visual grounding, enabling fine-grained cross-modal alignment between textual queries and image regions. However, transferring…

Domain GeneralizationSpatial ReasoningVisual Grounding

SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing

2025-12-09 · Aysim Toker, Andreea-Maria Oncescu, Roy Miles, Ismail Elezi 외 arxiv

Vision-language models (VLMs) are emerging as powerful generalist tools for remote sensing, capable of integrating information across diverse tasks and enabling flexible, instruction-based interactions via a chat interfa…

Spatial ReasoningVisual Grounding

TinyRS-R1: Compact Multimodal Language Model for Remote Sensing

2025-05-17 · Aybora Koksal, A. Aydin Alatan

Remote-sensing applications often run on edge hardware that cannot host today's 7B-parameter multimodal language models. This paper introduces TinyRS, the first 2B-parameter multimodal small language model (MSLM) optimiz…

Language ModelingLanguage ModellingOpen-Ended Question AnsweringQuestion Answering+4