paper-with-me

Papers

STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation

2025-06-16 · Jiamin Wang, Yichen Yao, Xiang Feng, Hang Wu, Yaming Wang, Qingqiu Huang, Yuexin Ma, Xinge Zhu

The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation and feature misalignment due to inadequate decoupling of spatio-temporal dynamics and limited cross-frame feature propagation mechanisms. To address these limitations, we present STAGE (Streaming Temporal Attention Generative Engine), a novel auto-regressive framework that pioneers hierarchical feature coordination and multi-phase optimization for sustainable video synthesis. To achieve high-quality long-horizon driving video generation, we introduce Hierarchical Temporal Feature Transfer (HTFT) and a novel multi-stage training strategy. HTFT enhances temporal consistency between video frames throughout the video generation process by modeling the temporal and denoising process separately and transferring denoising features between frames. The multi-stage training strategy is to divide the training into three stages, through model decoupling and auto-regressive inference process simulation, thereby accelerating model convergence and reducing error accumulation. Experiments on the Nuscenes dataset show that STAGE has significantly surpassed existing methods in the long-horizon driving video generation task. In addition, we also explored STAGE's ability to generate unlimited-length driving videos. We generated 600 frames of high-quality driving videos on the Nuscenes dataset, which far exceeds the maximum length achievable by existing methods.

📄 PDF Abstract BibTeX arXiv:2506.13138

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingDenoisingVideo Generation

Similar Papers 제목 키워드 기반

RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation

2025-09-18 · Yuming Jiang, Siteng Huang, Shengke Xue, Yaxi Zhao 외 arxiv

This paper presents RynnVLA-001, a vision-language-action(VLA) model built upon large-scale video generative pretraining from human demonstrations. We propose a novel two-stage pretraining methodology. The first stage, E…

Robot Manipulation

DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving

2026-04-01 · Yiyao Zhu, Ying Xue, Haiming Zhang, Guangfeng Jiang 외 arxiv

Vision-based autonomous driving has gained much attention due to its low costs and excellent performance. Compared with dense BEV (Bird's Eye View) or sparse query models, Gaussian-centric method is a comprehensive yet s…

Autonomous DrivingMotion Planning

Multi-Object Representation Learning via Feature Connectivity and Object-Centric Regularization

2023-09-21 · NeurIPS 2023 11

Discovering object-centric representations from images has the potential to greatly improve the robustness, sample efficiency and interpretability of machine learning algorithms. Current works on multi-object images typi…

Memory Storyboard: Leveraging Temporal Segmentation for Streaming Self-Supervised Learning from Egocentric Videos

2025-01-21 · Yanlai Yang, Mengye Ren

Self-supervised learning holds the promise to learn good representations from real-world continuous uncurated data streams. However, most existing works in visual self-supervised learning focus on static images or artifi…

Continual LearningContrastive LearningEvent SegmentationSelf-Supervised Learning

EgoForce: Robust Online Egocentric Motion Reconstruction via Diffusion Forcing

2026-05-13 · Inwoo Hwang, Donggeun Lim, Hojun Jang, Young Min Kim arxiv

With recent advances in embodied agents and AR devices, egocentric observations are readily available as input for real-world interactive online applications. However, egocentric viewpoints can only sporadically observe …