paper-with-me

홈 › Papers

WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception

2025-08-21 · Zhiheng Liu, Xueqing Deng, Shoufa Chen, Angtian Wang, Qiushan Guo, Mingfei Han, Zeyue Xue, Mengzhao Chen, Ping Luo, Linjie Yang arxiv

Generative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object structure and motion over extended durations. To address these issues, we introduce WorldWeaver, a robust framework for long video generation that jointly models RGB frames and perceptual conditions within a unified long-horizon modeling scheme. Our training framework offers three key advantages. First, by jointly predicting perceptual conditions and color information from a unified representation, it significantly enhances temporal consistency and motion dynamics. Second, by leveraging depth cues, which we observe to be more resistant to drift than RGB, we construct a memory bank that preserves clearer contextual information, improving quality in long-horizon video generation. Third, we employ segmented noise scheduling for training prediction groups, which further mitigates drift and reduces computational cost. Extensive experiments on both diffusion- and rectified flow-based models demonstrate the effectiveness of WorldWeaver in reducing temporal drift and improving the fidelity of generated videos.

📄 PDF Abstract BibTeX arXiv:2508.15720

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

2026-08-24 · Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li 외 arxiv

Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We presen…

Video Generation

Lyra 2.0: Explorable Generative 3D Worlds

2026-04-14 · Tianchang Shen, Sherwin Bahmani, Kai He, Sangeetha Grama Srinivasan 외 arxiv

Recent advances in video generation enable a new paradigm for 3D scene creation: generating camera-controlled videos that simulate scene walkthroughs, then lifting them to 3D via feed-forward reconstruction techniques. T…

Video Generation

SphericalDreamer: Generating Navigable Immersive 3D Worlds with Panorama Fusion

2026-05-19 · Antoine Schnepf, Karim Kassab, Flavian Vasile, Andrew Comport arxiv

The generation of immersive and navigable 3D environments is increasingly prevalent with the growing adoption of virtual reality and 3D content. However, recent methods face a fundamental limitation: they cannot produce …

A Recipe for Generating 3D Worlds From a Single Image

2025-03-20 · Katja Schwarz, Denys Rozumnyi, Samuel Rota Bulò, Lorenzo Porzi 외

We introduce a recipe for generating immersive 3D worlds from a single image by framing the task as an in-context learning problem for 2D inpainting models. This approach requires minimal training and uses existing gener…

In-Context Learning

AlayaWorld: Long-Horizon and Playable Video World Generation

2026-07-07 · AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan 외 arxiv

Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after deployment. Recent advances in video world …