paper-with-me

홈 › Papers

Perception-Aware Multimodal Spatial Reasoning from Monocular Images

2026-03-07 · Yanchun Cheng, Rundong Wang, Xulei Yang, Alok Prakash, Daniela Rus, Marcelo H Ang, ShiJie Li arxiv

Spatial reasoning from monocular images is essential for autonomous driving, yet current Vision-Language Models (VLMs) still struggle with fine-grained geometric perception, particularly under large scale variation and ambiguous object appearance. We propose a simple yet effective perception-aware multimodal reasoning framework that equips VLMs with explicit object-centric grounding ability. Instead of relying on textual bounding-box outputs, each referred object is represented using all Visual Reference Tokens (VRTs) within its spatial extent, enabling visual evidence and textual reasoning to be processed jointly in a unified token space. To further strengthen cross-modal interaction, we construct a Multimodal Chain-of-Thought (MM-CoT) dataset that injects aligned visual and textual reasoning signals. A deterministic ordering strategy is introduced to make supervision over inherently unordered VRT sets fully compatible with the VLM's autoregressive next-token prediction. With only standard supervised fine-tuning, our method achieves substantial improvements on the SURDS benchmark, outperforming previous approaches - including those using RL-based post-training - by a large margin across both single-object and multi-object tasks. These results demonstrate that accurate perception and multimodal reasoning are mutually reinforcing, and together form the key to robust spatial understanding in challenging monocular driving scenarios.

📄 PDF Abstract BibTeX arXiv:2603.06985

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningAutonomous DrivingSpatial Reasoning

Similar Papers 제목 키워드 기반

Spatial RoboGrasp: Generalized Robotic Grasping Control Policy

2025-05-27 · Yiqi Huang, Travis Davies, Jiahuan Yan, Jiankai Sun 외

Achieving generalizable and precise robotic manipulation across diverse environments remains a critical challenge, largely due to limitations in spatial perception. While prior imitation-learning approaches have made pro…

Depth EstimationImitation LearningMonocular Depth EstimationRobotic Grasping

SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving

2026-01-24 · Ashutosh Bajpai, Akshat Bhandari, Akshay Nambi, Tanmoy Chakraborty arxiv

Multimodal Small-to-Medium sized Language Models (MSLMs) have demonstrated strong capabilities in integrating visual and textual information but still face significant limitations in visual comprehension and mathematical…

Mathematical ReasoningData Augmentation

Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models

2026-03-18 · Kevin Qu, Haozhe Qi, Mihai Dusmanu, Mahdi Rad 외 arxiv

Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware reasoning. Recent efforts aim to augment th…

SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models

2025-11-10 · S Sakshi, Vaibhavi Lokegaonkar, Neil Zhang, Ramani Duraiswami 외 arxiv

Spatial perception is central to auditory intelligence, enabling accurate understanding of real-world acoustic scenes and advancing human-level perception of the world around us. While recent large audio-language models …

Spatial Reasoning

ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way

2025-07-11 · Rajarshi Roy, Devleena Das, Ankesh Banerjee, Arjya Bhattacharjee 외

We introduce ByDeWay, a training-free framework designed to enhance the performance of Multimodal Large Language Models (MLLMs). ByDeWay uses a novel prompting strategy called Layered-Depth-Based Prompting (LDP), which i…

Depth EstimationHallucinationLanguage ModelingLanguage Modelling+2