paper-with-me

Papers

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models

2025-05-22 · Runsen Xu, Weiyao Wang, Hao Tang, Xingyu Chen, Xiaodong Wang, Fu-Jen Chu, Dahua Lin, Matt Feiszli, Kevin J. Liang

Multi-modal large language models (MLLMs) have rapidly advanced in visual tasks, yet their spatial understanding remains limited to single images, leaving them ill-suited for robotics and other real-world applications that require multi-frame reasoning. In this paper, we propose a framework to equip MLLMs with robust multi-frame spatial understanding by integrating depth perception, visual correspondence, and dynamic perception. Central to our approach is the MultiSPA dataset, a novel, large-scale collection of more than 27 million samples spanning diverse 3D and 4D scenes. Alongside MultiSPA, we introduce a comprehensive benchmark that tests a wide spectrum of spatial tasks under uniform metrics. Our resulting model, Multi-SpatialMLLM, achieves significant gains over baselines and proprietary systems, demonstrating scalable, generalizable multi-frame reasoning. We further observe multi-task benefits and early indications of emergent capabilities in challenging scenarios, and showcase how our model can serve as a multi-frame reward annotator for robotics.

📄 PDF Abstract BibTeX arXiv:2505.17015

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models

2026-03-26 · Huizhi Liang, Yichao Shen, Yu Deng, Sicheng Xu 외 arxiv

Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high-level spatia…

Spatial Reasoning

CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence

2026-06-09 · Yiming Zhang, Ruoxuan Cao, Zhihang Zhong arxiv

Spatial intelligence is a key frontier for multimodal large language models (MLLMs), enabling them to reason about the physical world from visual experience. Inspired by human spatial cognition, recent approaches constru…

Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning

2025-10-03 · Chashi Mahiul Islam, Oteo Mamo, Samuel Jacob Chacko, Xiuwen Liu 외 arxiv

Vision-language models (VLMs) have advanced multimodal reasoning but still face challenges in spatial reasoning for 3D scenes and complex object configurations. To address this, we introduce SpatialViLT, an enhanced VLM …

Multimodal ReasoningMulti-Task LearningSpatial Reasoning

Spatial Audio Motion Understanding and Reasoning

2025-09-18 · Arvind Krishna Sridhar, Yinyi Guo, Erik Visser arxiv

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving…

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding

2026-07-16 · Xiao Lin, Xiaohu Huang, Kai Han arxiv

Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhan…

Spatial Reasoning