paper-with-me

홈 › Papers

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

2025-07-06 · Binbin Ji, Siddharth Agrawal, Qiance Tang, Yvonne Wu arxiv

This study investigates the spatial reasoning capabilities of vision-language models (VLMs) through Chain-of-Thought (CoT) prompting and reinforcement learning. We begin by evaluating the impact of different prompting strategies and find that simple CoT formats, where the model generates a reasoning step before the answer, not only fail to help, but can even harm the model's original performance. In contrast, structured multi-stage prompting based on scene graphs (SceneGraph CoT) significantly improves spatial reasoning accuracy. Furthermore, to improve spatial reasoning ability, we fine-tune models using Group Relative Policy Optimization (GRPO) on the SAT dataset and evaluate their performance on CVBench. Compared to supervised fine-tuning (SFT), GRPO achieves higher accuracy on Pass@1 evaluations and demonstrates superior robustness under out-of-distribution (OOD) conditions. In particular, we find that SFT overfits to surface-level linguistic patterns and may degrade performance when test-time phrasing changes (e.g., from "closer to" to "farther from"). GRPO, on the other hand, generalizes more reliably and maintains stable performance under such shifts. Our findings provide insights into how reinforcement learning and structured prompting improve the spatial reasoning capabilities and generalization behavior of modern VLMs. All code is open source at: https://github.com/Yvonne511/spatial-vlm-investigator

📄 PDF Abstract BibTeX arXiv:2507.13362

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningSpatial Reasoning

Similar Papers 제목 키워드 기반

Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation

2025-01-01 · CVPR 2025 1 · Yudi Shi, Shangzhe Di, Qirui Chen, Weidi Xie

This paper tackles the problem of video question answering (VideoQA), a task that often requires multi-step reasoning and a profound understanding of spatial-temporal dynamics. While large video-language models perfo…

Language ModelingLanguage ModellingLarge Language ModelMultiple-choice+2

SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning

2026-03-23 · Byungwoo Jeon, Dongyoung Kim, Huiwon Jang, Insoo Kim 외 arxiv

Despite the remarkable success of large-scale pre-trained image representation models (i.e., vision encoders) across various vision tasks, they are predominantly trained on 2D image data and therefore often fail to captu…

SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

2025-01-17 · Yuecheng Liu, Dafeng Chi, Shiguang Wu, Zhanguang Zhang 외

Spatial reasoning is an essential problem in embodied AI research. Efforts to enhance spatial reasoning abilities through supplementary spatial data and fine-tuning have proven limited and ineffective when addressing com…

Spatial ReasoningTask Planning

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

2026-06-16 · Yatai Ji, An-Chieh Cheng, Yang Fu, Yukang Chen 외 arxiv

Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging. Moreover, different spatial…

Reinforcement LearningSpatial Reasoning

Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning

2026-03-21 · Xueqi Ma, Shuo Yang, Yanbei Jiang, Shu Liu 외 arxiv

Despite remarkable advances in large Vision-Language Models (VLMs), spatial reasoning remains a persistent challenge. In this work, we investigate how attention heads within VLMs contribute to spatial reasoning by analyz…

Relational ReasoningSpatial Reasoning