paper-with-me

Papers

X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving

2026-03-20 · Chaoda Zheng, Sean Li, Jinhao Deng, Zhennan Wang, Shijia Chen, Liqiang Xiao, Ziheng Chi, Hongbin Lin, Kangjie Chen, Boyang Wang, Yu Zhang, Xianming Liu arxiv

Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still rely heavily on real-world road testing, which is costly, biased toward limited scenario coverage, and difficult to reproduce. These challenges motivate a real-world simulator that can generate realistic future observations under proposed actions, while remaining controllable and stable over long horizons. We present X-World, an action-conditioned multi-camera generative world model that simulates future observations directly in video space. Given synchronized multi-view camera history and a future action sequence, X-World generates future multi-camera video streams that follow the commanded actions. To ensure reproducible and editable scene rollouts, X-World further supports optional controls over dynamic traffic agents and static road elements, and retains a text-prompt interface for appearance-level control (e.g., weather and time of day). Beyond world simulation, X-World also enables video style transfer by conditioning on appearance prompts while preserving the underlying action and scene dynamics. At the core of X-World is a multi-view latent video generator designed to explicitly encourage cross-view geometric consistency and temporal coherence under diverse control signals. Experiments show that X-World achieves high-quality multi-view video generation with (i) strong view consistency across cameras, (ii) stable temporal dynamics over long rollouts, and (iii) high controllability with strict action following and faithful adherence to optional scene controls. These properties make X-World a practical foundation for scalable and reproducible evaluation.

📄 PDF Abstract BibTeX arXiv:2603.19979

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingVideo GenerationStyle Transfer

Similar Papers 제목 키워드 기반

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

2026-06-26 · Haoyuan Wang, Yabo Chen, Haibin Huang, Chi Zhang 외 arxiv

Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video generation offers a scalable foundation, but suffers from error accu…

Video Generation

HumanVid: Demystifying Training Data for Camera-controllable Human Image Animation

2024-07-24 · Zhenzhi Wang, Yixuan Li, Yanhong Zeng, Youqing Fang 외

Human image animation involves generating videos from a character photo, allowing user control and unlocking the potential for video and movie production. While recent approaches yield impressive results using high-quali…

BenchmarkingHuman AnimationImage AnimationVideo Generation

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

2026-07-30 · Zexuan Yan, Yuzhou Wu, Yue Ma, Zonghang He 외 arxiv

Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action…

Video Generation

Prisma-World: Camera-Controllable Multi-Agent Video World Model

2026-06-08 · Huiqiang Sun, Zhan Peng, Size Wu, Kun Wang 외 arxiv

Video world models have made rapid progress in generating controllable visual experiences, but most of them still simulate the world from a single observer. Extending such models to multiple agents raises a central chall…

EgoCS-400K: An Egocentric Gameplay Dataset for World Models

2026-06-16 · Rongjin Guo, Dong Liang, Yuhao Liu, Fang Liu 외 arxiv

The shift from video generation to interactive world modeling places new demands on data: beyond captioned videos, world models require temporally aligned video-action-language trajectories grounded in the actions, camer…

Action UnderstandingVideo Generation