paper-with-me

Papers

SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery

2025-12-08 · Meng Cao, Xingyu Li, Xue Liu, Ian Reid, Xiaodan Liang arxiv

Despite advancements in Multi-modal Large Language Models (MLLMs) for scene understanding, their performance on complex spatial reasoning tasks requiring mental simulation remains significantly limited. Current methods often rely on passive observation of spatial data, failing to internalize an active mental imagery process. To bridge this gap, we propose SpatialDreamer, a reinforcement learning framework that enables spatial reasoning through a closedloop process of active exploration, visual imagination via a world model, and evidence-grounded reasoning. To address the lack of fine-grained reward supervision in longhorizontal reasoning tasks, we propose Geometric Policy Optimization (GeoPO), which introduces tree-structured sampling and step-level reward estimation with geometric consistency constraints. Extensive experiments demonstrate that SpatialDreamer delivers highly competitive results across multiple challenging benchmarks, signifying a critical advancement in human-like active spatial mental simulation for MLLMs.

📄 PDF Abstract BibTeX arXiv:2512.07733

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningScene UnderstandingSpatial Reasoning

Similar Papers 제목 키워드 기반

SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input

2024-11-18 · CVPR 2025 1 · Zhen Lv, Yangqi Long, Congzhentao Huang, Cao Li 외

Stereo video synthesis from a monocular input is a demanding task in the fields of spatial computing and virtual reality. The main challenges of this task lie on the insufficiency of high-quality paired stereo videos for…

Novel View SynthesisVideo Generation

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation

2026-05-15 · Yiming Zhao, Yu Zeng, Wenxuan Huang, Zhen Fang 외 arxiv

Large Vision-Language Models (LVLMs) have shown significant progress in video understanding, yet they face substantial challenges in tasks requiring precise spatiotemporal localization at the instance level. Existing met…

ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning

2025-12-02 · Yifan Li, Yingda Yin, Lingting Zhu, Weikai Chen 외 arxiv

Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances. Yet existing solutions generally collap…

Video Object SegmentationReinforcement LearningVideo Segmentation

GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes

2025-11-27 · Di Wang, Shunyu Liu, Wentao Jiang, Fengxiang Wang 외 arxiv

Multimodal large language models (MLLMs) have undergone rapid development in advancing geospatial scene understanding. Recent studies have sought to enhance the reasoning capabilities of remote sensing MLLMs, typically t…

Reinforcement LearningScene Understanding

Incentivizing Reasoning from Weak Supervision

2025-05-26 · Yige Yuan, Teng Xiao, Shuchang Tao, Xue Wang 외

Large language models (LLMs) have demonstrated impressive performance on reasoning-intensive tasks, but enhancing their reasoning abilities typically relies on either reinforcement learning (RL) with verifiable signals o…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)