paper-with-me

홈 › Papers

Unified 3D Scene Understanding Through Physical World Modeling

2026-05-23 · Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh, Jared Watrous, Honglin Chen, Khai Loong Aw, Daniel L. K. Yamins arxiv

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essential for perception and interaction. Existing approaches have typically addressed these tasks in isolation, preventing them from sharing a common representation or transferring knowledge across tasks. A conceptually simpler but practically non-trivial alternative is to unify these diverse tasks into a single model, reducing different tasks from separate training objectives to merely different prompts and allowing for joint training across all datasets. In this work, we present a physical world model for unified 3D understanding and interaction (3WM), formulated as a probabilistic graphical model in which nodes represent multimodal scene elements such as RGB, optical flow, and camera pose. Diverse tasks emerge from different inference pathways through the graph: novel view synthesis from RGB and dense flow prompts, object manipulation from RGB and sparse flow prompts, and depth estimation from RGB and camera conditioning, all zero-shot without task-specific training. 3WM outperforms specialized baselines without the need for finetuning by offering precise controllability, strong geometric consistency, and robustness in real-world scenarios, achieving state-of-the-art performance on NVS and 3D object manipulation. Beyond predefined tasks, the model supports composable inference pathways, such as moving objects aside while navigating a 3D environment, enabling complex geometric reasoning. This demonstrates that a unified model can serve as a practical alternative to fragmented task-specific systems, taking a step towards a general-purpose visual world model.

📄 PDF Abstract BibTeX arXiv:2605.24321

Code (0)

등록된 구현이 없습니다.

Tasks

Novel View SynthesisScene UnderstandingDepth EstimationVisual Reasoning

Similar Papers 제목 키워드 기반

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

2026-04-30 · Xin Zhou, Dingkang Liang, Xiwu Chen, Feiyang Tan 외 arxiv

Driving world models serve as a pivotal technology for autonomous driving by simulating environmental dynamics. However, existing approaches predominantly focus on future scene generation, often overlooking comprehensive…

Scene UnderstandingAutonomous DrivingScene Generation

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

2025-01-24 · Xin Zhou, Dingkang Liang, Sifan Tu, Xiwu Chen 외

Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to incorporate scene understanding, which invo…

Autonomous DrivingLanguage ModelingLanguage ModellingLarge Language Model+3

Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation

2025-09-28 · Hanyu Zhou, Gim Hee Lee arxiv

Vision-language models (VLMs) have demonstrated strong performance in 2D scene understanding and generation, but extending this unification to the physical world remains an open challenge. Existing 3D and 4D approaches t…

Scene Understanding

HoloScene: Simulation-Ready Interactive 3D Worlds from a Single Video

2025-10-07 · Hongchi Xia, Chih-Hao Lin, Hao-Yu Hsu, Quentin Leboutet 외 arxiv

Digitizing the physical world into accurate simulation-ready virtual environments offers significant opportunities in a variety of fields such as augmented and virtual reality, gaming, and robotics. However, current 3D r…

3D Reconstruction

Prediction of Scene Plausibility

2022-12-02 · Or Nachmias, Ohad Fried, Ariel Shamir

Understanding the 3D world from 2D images involves more than detection and segmentation of the objects within the scene. It also includes the interpretation of the structure and arrangement of the scene elements. Such un…

PredictionScene Understanding