paper-with-me

홈 › Papers

SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models

2025-10-09 · Hongxing Li, Dingming Li, Zixuan Wang, Yuchen Yan, Hang Wu, Wenqi Zhang, Yongliang Shen, Weiming Lu, Jun Xiao, Yueting Zhuang arxiv

Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We identify that this limitation stems from a critical gap: existing methods attempt to learn spatial reasoning directly without establishing the hierarchical foundations of perception and understanding. To address this challenge, we present a comprehensive methodology for building spatial intelligence progressively. We introduce SpatialLadder-26k, a multimodal dataset containing 26,610 samples spanning object localization, single image, multi-view, and video spatial reasoning tasks, constructed through a standardized pipeline that ensures systematic coverage across modalities. Building on this dataset, we design a three-stage progressive training framework that (1) establishes spatial perception through object localization, (2) develops spatial understanding through multi-dimensional spatial tasks, and (3) strengthens complex reasoning via reinforcement learning with verifiable rewards. This approach yields SpatialLadder, a 3B-parameter model that achieves state-of-the-art performance on spatial reasoning benchmarks, with 23.4% average improvement over the base model, surpassing GPT-4o by 20.8% and Gemini-2.0-Flash by 10.1%. Notably, SpatialLadder maintains strong generalization with 7.2% improvement on out-of-domain benchmarks, demonstrating that progressive training from perception to reasoning is essential for robust spatial intelligence.

📄 PDF Abstract BibTeX arXiv:2510.08531

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningObject LocalizationSpatial Reasoning

Similar Papers 제목 키워드 기반

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs

2026-04-01 · Zhenghao Chen, Huiqun Wang, Di Huang arxiv

Multimodal large language models (MLLMs) are increasingly being applied to spatial cognition tasks, where they are expected to understand and interact with complex environments. Most existing works improve spatial reason…

Spatial Reasoning

SpaceVLN: A Zero-Shot Vision-and-Language Navigation Agent with Online Spatial Cognitive Memory and Reasoning

2026-06-08 · Yucheng Deng, Pingrui Lai, Xinhai Li, Chenjia Bai 외 arxiv

Vision-and-Language Navigation in continuous environments requires agents to understand the spatial structure of previously unseen environments in order to follow language instructions. Although foundation models have op…

Spatial Reasoning

SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning

2026-03-23 · Byungwoo Jeon, Dongyoung Kim, Huiwon Jang, Insoo Kim 외 arxiv

Despite the remarkable success of large-scale pre-trained image representation models (i.e., vision encoders) across various vision tasks, they are predominantly trained on 2D image data and therefore often fail to captu…

PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly

2025-06-10 · Liang Ma, Jiajun Wen, Min Lin, Rongtao Xu 외

While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, …

Question AnsweringScene UnderstandingSpatial ReasoningVisual Question Answering+1

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

2026-03-28 · Jian Zhang, Shijie Zhou, Bangya Liu, Achuta Kadambi 외 arxiv

Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This limitation arises from their inability to capture fine-grained 3D geomet…

Spatial Reasoning