paper-with-me

홈 › Papers

GeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement Learning

2026-01-07 · Wenshuai Li, Xiantai Xiang, Zixiao Wen, Guangyao Zhou, Ben Niu, Feng Wang, Lijia Huang, Qiantong Wang, Yuxin Hu arxiv

The evolution of Remote Sensing Vision-Language Models(RS-VLMs) emphasizes the importance of transitioning from perception-centric recognition toward high-level deductive reasoning to enhance cognitive reliability in complex spatial tasks. However, current models often suffer from logical hallucinations, where correct answers are derived from flawed reasoning chains or rely on positional shortcuts rather than spatial logic. This decoupling undermines reliability in strategic spatial decision-making. To address this, we present GeoReason, a framework designed to synchronize internal thinking with final decisions. We first construct GeoReason-Bench, a logic-driven dataset containing 4,000 reasoning trajectories synthesized from geometric primitives and expert knowledge. We then formulate a two-stage training strategy: (1) Supervised Knowledge Initialization to equip the model with reasoning syntax and domain expertise, and (2) Consistency-Aware Reinforcement Learning to refine deductive reliability. This second stage integrates a novel Logical Consistency Reward, which penalizes logical drift via an option permutation strategy to anchor decisions in verifiable reasoning traces. Experimental results demonstrate that our framework significantly enhances the cognitive reliability and interpretability of RS-VLMs, achieving state-of-the-art performance compared to other advanced methods.

📄 PDF Abstract BibTeX arXiv:2601.04118

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Vision-Language Models in Remote Sensing: Current Progress and Future Trends

2023-05-09 · Xiang Li, Congcong Wen, Yuan Hu, Zhenghang Yuan 외

The remarkable achievements of ChatGPT and GPT-4 have sparked a wave of interest and research in the field of large language models for Artificial General Intelligence (AGI). These models provide intelligent solutions cl…

Image CaptioningImage GenerationImage Retrievalobject-detection+6

DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception

2025-07-30 · Pei Deng, Wenqian Zhou, Hanlin Wu arxiv

Accurate interpretation of land-cover changes in multi-temporal satellite imagery is critical for real-world scenarios. However, existing methods typically provide only one-shot change masks or static captions, limiting …

Visual Question AnsweringChange Detection

VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis

2024-03-29 · Chao Pang, Xingxing Weng, Jiang Wu, Jiayu Li 외

This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest…

HallucinationImage CaptioningLanguage ModelingLanguage Modelling+6

GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model

2024-06-03 · Ling Li, Yu Ye, Bingchuan Jiang, Wei Zeng

This work tackles the problem of geo-localization with a new paradigm using a large vision-language model (LVLM) augmented with human inference knowledge. A primary challenge here is the scarcity of data for training the…

geo-localizationLanguage ModelingLanguage Modelling

RSVQA: Visual Question Answering for Remote Sensing Data

2020-03-16 · Sylvain Lobry, Diego Marcos, Jesse Murray, Devis Tuia

This paper introduces the task of visual question answering for remote sensing data (RSVQA). Remote sensing images contain a wealth of information which can be useful for a wide range of tasks including land cover classi…

Land Cover ClassificationObject CountingQuestion AnsweringVisual Question Answering+1