paper-with-me

홈 › Papers

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

2026-06-25 · Bin Hu, Yanwen Ma, Jiehui Huang, Ziliang Zhang, Haoning Wu, Ruicheng Zhang, Yaokun Li, Zijun Wang, Yuechen Zhang, Chun-Mei Tseng, Hanhui Li, Shengju Qian, Jun Zhou, Kaipeng Zhang, Xiaodan Liang, Jiaya Jia, Xiu Li arxiv

Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering trajectories, and physical dynamics are typically learned as implicit correlations from data rather than as controllable variables. This limitation hinders their applicability to authored game environments, where physical rules are deliberately designed and require explicit manipulation. We introduce PhysEditWorld, a multimodal dataset with physical parameters, with a primary focus on gravity in this initial version. At its core, PhysEditWorld is built upon a replay paradigm implemented with a UE5 replay-and-rendering pipeline. Each scenario records a normalized action trace and replays the same initial state, character controller, action sequence, and camera policy under multiple gravity configurations, enabling controlled and attributable physical variation. PhysEditWorld contains 12 cinematic UE5 scenes, over 100 hours of gameplay interactions, and more than 60 million rendered rollout frames. Each sample provides synchronized multimodal signals, including RGB, depth, normals, audio, action traces, camera trajectory, engine states, semantic annotations, and explicit gravity labels. We further conduct initial utility studies on both generative video models and world understanding models, demonstrating that PhysEditWorld enables improved gravity-faithful dynamics modeling, enhances consistency under physical edits, and provides a scalable foundation for controllable world modeling research.

📄 PDF Abstract BibTeX arXiv:2606.26694

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Real2Sim: A Physics-driven and Editable Gaussian Splatting Framework for Autonomous Driving Scenes

2026-05-13 · Kaicong Huang, Talha Azfar, Weisong Shi, Ruimin Ke arxiv

Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and traditional simulation suffers from a persistent reality gap. While recen…

Trajectory PredictionAutonomous Driving

DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces

2026-03-05 · Mohammad Sadil Khan, Muhammad Usama, Rolandos Alexandros Potamias, Didier Stricker 외 arxiv

Computer-Aided Design (CAD) relies on structured and editable geometric representations, yet existing generative methods are constrained by small annotated datasets with explicit design histories or boundary representati…

4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation

2026-08-27 · Zehao Qi, Haochen Luo, Jia-Wang Bian, Zeyu Ma 외 arxiv

Embodied agents need environments that are visually diverse, physically interactive, and changing over time. Procedural simulators can generate large interactive scene collections, and recent 4D generators produce compel…

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation

2026-07-02 · Chunjiang Liu, Xiaoyuan Wang, Haoyu Chen, Yizhou Zhao 외 arxiv

LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, in which liquids flow, particles emit, ri…

Video Generation

FalconTrack: Photorealistic Auto-Labeled Perception and Physics-Aware Vision-Based Aerial Tracking

2026-06-29 · Yan Miao, Karteek Gandiboyina, Noah Giles, Hideki Okamoto 외 arxiv

Vision-based aerial tracking is critical in GPS-denied environments. Reliable perception for tracking depends on large-scale labeled data, yet most photorealistic datasets rely on heavy manual annotation and are time-con…

Visual Tracking