paper-with-me

Papers

Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

2026-03-30 · Muyang He, Hanzhong Guo, Junxiong Lin, Yizhou Yu arxiv

The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. However, a critical gap still remains between the theoretical capacity for world simulation and the heavy computational costs of spatiotemporal modeling. To address this, we comprehensively and systematically review video generation frameworks and techniques that consider efficiency as a crucial requirement for practical world modeling. We introduce a novel taxonomy in three dimensions: efficient modeling paradigms, efficient network architectures, and efficient inference algorithms. We further show that bridging this efficiency gap directly empowers interactive applications such as autonomous driving, embodied AI, and game simulation. Finally, we identify emerging research frontiers in efficient video-based world modeling, arguing that efficiency is a fundamental prerequisite for evolving video generators into general-purpose, real-time, and robust world simulators. A curated GitHub repository of the reviewed literature can be found at https://github.com/Isaachhh/Efficient-VWM-Survey.

📄 PDF Abstract BibTeX arXiv:2603.28489

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingVideo Generation

Similar Papers 제목 키워드 기반

A Mechanistic View on Video Generation as World Models: State and Dynamics

2026-01-22 · Luozhou Wang, Zhifei Chen, Yihua Du, Dongyu Yan 외 arxiv

Large-scale video generation models have demonstrated emergent physical coherence, positioning them as potential world models. However, a gap remains between contemporary "stateless" video architectures and classic state…

Video Generation

UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation

2025-12-08 · Jiehui Huang, Yuechen Zhang, Xu He, Yuan Gao 외 arxiv

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal i…

Zero-shot GeneralizationMulti-Task LearningVideo Generation

World Model for Robot Learning: A Comprehensive Survey

2026-04-30 · Bohan Hou, Gen Li, Jindou Jia, Tuo An 외 arxiv

World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They support policy learning, planning, simulation, evaluation, data generat…

Reinforcement LearningAutonomous DrivingVideo Generation

MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation

2026-02-27 · Haoyuan Shi, Yunxin Li, Nanhao Deng, Zhenran Xu 외 arxiv

The evolution of video generation toward complex, multi-shot narratives has exposed a critical deficit in current evaluation methods. Existing benchmarks remain anchored to single-shot paradigms, lacking the comprehensiv…

Video Generation

Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models

2026-03-23 · Meiqi Wu, Zhixin Cai, Fufangchen Zhao, Xiaokun Feng 외 arxiv

Video--based world models have emerged along two dominant paradigms: video generation and 3D reconstruction. However, existing evaluation benchmarks either focus narrowly on visual fidelity and text--video alignment for …

3D ReconstructionVideo GenerationVideo Alignment