paper-with-me

홈 › Papers

AirScape: An Aerial Generative World Model with Motion Controllability

2025-07-10 · Baining Zhao, Rongze Tang, Mingyuan Jia, Ziyou Wang, Fanghang Man, Xin Zhang, Yu Shang, Weichen Zhang, Wei Wu, Chen Gao, Xinlei Chen, Yong Li arxiv

How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general spatial imagination capability, we present AirScape, the first world model designed for six-degree-of-freedom aerial agents. AirScape predicts future observation sequences based on current visual inputs and motion intentions. Specifically, we construct a dataset for aerial world model training and testing, which consists of 11k video-intention pairs. This dataset includes first-person-view videos capturing diverse drone actions across a wide range of scenarios, with over 1,000 hours spent annotating the corresponding motion intentions. Then we develop a two-phase schedule to train a foundation model--initially devoid of embodied spatial knowledge--into a world model that is controllable by motion intentions and adheres to physical spatio-temporal constraints. Experimental results demonstrate that AirScape significantly outperforms existing foundation models in 3D spatial imagination capabilities, especially with over a 50% improvement in metrics reflecting motion alignment. The project is available at: https://embodiedcity.github.io/AirScape/.

📄 PDF Abstract BibTeX arXiv:2507.08885

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls

2026-05-19 · Abdul Mohaimen Al Radi, Kunyang Li, Yuzhang Shang, Mubarak Shah 외 arxiv

Foundation video models produce visually impressive results, but their use in embodied AI remains limited because they are primarily trained on natural language rather than low-level control signals. This limitation is e…

Video Generation

Nonlinear Controllability Assessment of Aerial Manipulator Systems using Lagrangian Reduction

2021-08-15 · Skylar X. Wei, Matthew R. Burkhardt, Joel Burdick

This paper analyzes the nonlinear Small-Time Local Controllability (STLC) of a class of underatuated aerial manipulator robots. We apply methods of Lagrangian reduction to obtain their lowest dimensional equations of mot…

PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth

2025-05-03 · Bu Jin, Weize Li, Baihan Yang, Zhenxin Zhu 외

Recent advancements in autonomous driving (AD) systems have highlighted the potential of world models in achieving robust and generalizable performance across both ordinary and challenging driving conditions. However, a …

Autonomous DrivingCamera Pose EstimationDepth EstimationPose Estimation+1

Dreamland: Controllable World Creation with Simulator and Generative Models

2025-06-09 · Sicheng Mo, Ziyang Leng, Leon Liu, Weizhen Wang 외

Large-scale video generative models can synthesize diverse and realistic visual content for dynamic world creation, but they often lack element-wise controllability, hindering their use in editing scenes and training emb…

OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction

2025-09-04 · Bu Jin, Songen Gu, Xiaotao Hu, Yupeng Zheng 외 arxiv

In this paper, we propose OccTENS, a generative occupancy world model that enables controllable, high-fidelity long-term occupancy generation while maintaining computational efficiency. Different from visual generation, …

Computational Efficiency