paper-with-me

Papers

Thinking with Spatial Code for Physical-World Video Reasoning

2026-03-05 · Jieneng Chen, Wenxin Ma, Ruisheng Yuan, Yunzhi Zhang, Jiajun Wu, Alan Yuille arxiv

We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical-world visual question answering. We highlight the empirical finding that our proposed spatial encoder can parse videos into structured spatial code with explicit 3D oriented bounding boxes and semantic labels, enabling large language models (LLMs) to reason directly over explicit spatial variables. Specifically, we propose the spatial encoder that encodes image and geometric features by unifying 6D object parsing and tracking backbones with geometric prediction, and we further finetuning LLMs with reinforcement learning using a spatial rubric reward that encourages perspective-aware, geometrically grounded inference. As a result, our model outperforms proprietary vision-language models on VSI-Bench, setting a new state-of-the-art. Code is available at https://github.com/Beckschen/spatialcode.

📄 PDF Abstract BibTeX arXiv:2603.05591

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringReinforcement Learning

Similar Papers 제목 키워드 기반

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

2026-03-13 · Yuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu 외 arxiv

Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While current Multimodal Large Language Models (MLL…

Visual Question Answering

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

2026-05-07 · Zhangquan Chen, Manyuan Zhang, Xinlei Yu, Xiang An 외 arxiv

Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vision-language models (VLMs). Prior approaches either verbalize spatial-te…

Reinforcement LearningSpatial Reasoning

Visuospatial Cognitive Assistant

2025-05-18 · Qi Feng

Video-based spatial cognition is vital for robotics and embodied AI but challenges current Vision-Language Models (VLMs). This paper makes two key contributions. First, we introduce ViCA (Visuospatial Cognitive Assistant…

Spatial Reasoning

Teaching Video Diffusion Model with Latent Physical Phenomenon Knowledge

2024-11-18 · Qinglong Cao, Ding Wang, Xirui Li, Yuntian Chen 외

Video diffusion models have exhibited tremendous progress in various video generation tasks. However, existing models struggle to capture latent physical knowledge, failing to infer physical phenomena that are challengin…

Video Generation

MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators

2024-04-07 · Shenghai Yuan, Jinfa Huang, Yujun Shi, Yongqi Xu 외

Recent advances in Text-to-Video generation (T2V) have achieved remarkable success in synthesizing high-quality general videos from textual descriptions. A largely overlooked problem in T2V is that existing models have n…

Text-to-Video GenerationVideo Generation