paper-with-me

Papers

Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning

2025-11-20 · Yibin Huang, Wang Xu, Wanyue Zhang, Helu Zhi, Jingjing Huang, Yangbin Xu, Yangang Sun, Conghui Zhu, Tiejun Zhao arxiv

Spatial intelligence is a critical frontier for Multimodal Large Language Models (MLLMs), empowering them to comprehend the physical world. Drawing inspiration from human perception mechanisms, prior studies attempt to construct a spatial understanding via grid-based cognitive maps. However, current grid-based map methods rely on discretized representations, which limit the model's ability in fine-grained spatial reasoning. To overcome this limitation, we propose Video2Layout, a framework for reconstructing metric-grounded spatial layouts from video. The framework uses continuous object boundary coordinates to enable quantitative spatial computation, which effectively reduces ambiguity in natural language descriptions of spatial relationships. Specifically, our method comprises two stages. First, in supervised fine-tuning stage, we construct a high-quality dataset from the AI2THOR simulator, which enables the model to learn the mapping from visual inputs to precise boundary coordinates. Subsequently, a reinforcement fine-tuning stage enhances the model's real-world generalization capabilities. Based on the above framework, we investigate factors that affect cognitive map accuracy and quantify its relationship with task performance. Evaluated on mainstream spatial reasoning benchmarks, our model, V2LO-7B, achieves an average improvement of 3.24\% over the model trained on grid maps, validating the superiority of our method.

📄 PDF Abstract BibTeX arXiv:2511.16160

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

One Video, One World: Turning Monocular Video into Physical 4D Scenes

2026-06-30 · Junhao Chen, Boran Zhang, Mingjin Chen, Henghaofan Zhang 외 arxiv

We introduce \textbf{OVOW}, the first training-free system that reconstructs \emph{instance-level, simulation-ready} 4D mesh scenes from a single monocular video. Recent 4D reconstruction achieves impressive rendering qu…

Point Clouds

LLM-grounded Video Diffusion Models

2023-09-29 · Long Lian, Baifeng Shi, Adam Yala, Trevor Darrell 외

Text-conditioned diffusion models have emerged as a promising tool for neural video generation. However, current models still struggle with intricate spatiotemporal prompts and often generate restricted or incorrect moti…

Language ModelingLanguage ModellingLarge Language ModelVideo Generation

FloorSAM: SAM-Guided Floorplan Reconstruction with Semantic-Geometric Fusion

2025-09-19 · Han Ye, Haofu Wang, Yunchi Zhang, Jiangjian Xiao 외 arxiv

Reconstructing building floor plans from point cloud data is key for indoor navigation, BIM, and precise measurements. Traditional methods like geometric algorithms and Mask R-CNN-based deep learning often face issues wi…

Zero-Shot LearningImage Enhancement

ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation

2026-05-15 · Michał Ciesiółka, Dawid Wiśniewski, Adrian Charkiewicz, Kamil Guttmann arxiv

We present ForMaT (Format-Preserving Multilingual Translation), a parallel corpus of 3,956 PDFs across 15 language pairs that preserves original layout metadata proposed for multimodal machine translation. To ensure stru…

Multimodal Machine Translation

StructuredMesh: 3D Structured Optimization of Façade Components on Photogrammetric Mesh Models using Binary Integer Programming

2023-06-07 · Libin Wang, Han Hu, Qisen Shang, Bo Xu 외

The lack of fa\c{c}ade structures in photogrammetric mesh models renders them inadequate for meeting the demands of intricate applications. Moreover, these mesh models exhibit irregular surfaces with considerable geometr…

object-detectionObject Detection