paper-with-me

Papers

Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion

2026-08-27 · Bowen Xue, Brandon Y. Feng, Chenguo Lin, Yuchen Lin, Yujia Zeng, Lvmin Zhang, Maneesh Agrawala, Honglei Yan, Panwang Pan arxiv

Scaling video generation to long durations reveals a critical bottleneck: current models lack robust long-term memory. This deficiency can be studied along two critical aspects: object permanence, the ability to precisely reproduce the appearance of objects upon re-entry; and memory capacity, the ability to process ultra-long context and use information from distant history. Robust long-term memory requires both: object permanence without sufficient context handling limits the temporal scope, while long context length without permanence fails to maintain identity. To address this, we present Ring Forcing, an autoregressive video diffusion framework designed to robustly construct and precisely utilize long-term memory. Our ring-structured training strategy enforces retrieval from distant history, effectively reconciling the trade-off between strict historical adherence and generative diversity. To expand memory capacity, we introduce a compression and timestep composition strategy. Under fixed sequence length constraints, this method extends the effective historical span to minutes-long durations and achieves a comprehensive receptive field over the entire history. Furthermore, we present a sparse RoPE mechanism to enable flexible, scalable memory adaptation while fully exploiting pre-trained priors. Extensive experiments demonstrate that Ring Forcing achieves superior minutes-long coherence and object permanence, significantly outperforming state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2608.26794

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation

2026-03-22 · Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song 외 arxiv

Autoregressive video diffusion has recently emerged as a promising paradigm for long-video generation, enabling causal synthesis beyond the temporal limits of bidirectional models. Existing forcing-based training strateg…

Video Generation

Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation

2026-05-15 · Mingqiang Wu, Weilun Feng, Zhefeng Zhang, Haotong Qin 외 arxiv

Autoregressive video diffusion models enable open-ended generation through local attention and KV caching. However, existing training-free long-video optimization methods mainly focus on stable extension under a single p…

Video Generation

RELIC: Interactive Video World Model with Long-Horizon Memory

2025-12-03 · Yicong Hong, Yiqun Mei, Chongjian Ge, Yiran Xu 외 arxiv

A truly interactive world model requires three key ingredients: real-time long-horizon streaming, consistent spatial memory, and precise user control. However, most existing approaches address only one of these aspects i…

Wonder: Video World Model Done Better

2026-07-28 · Jiacong Xu, Hanwen Jiang, Zhixin Shu, Kalyan Sunkavalli 외 arxiv

We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactivel…

Video Generation

Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft

2025-10-03 · Junchao Huang, Xinting Hu, Boyao Han, Shaoshuai Shi 외 arxiv

Autoregressive video diffusion models have proved effective for world modeling and interactive scene generation, with Minecraft gameplay as a representative application. To faithfully simulate play, a model must generate…

Computational Efficiency3D ReconstructionScene Generation