paper-with-me

홈 › Papers

Context Forcing: Consistent Autoregressive Video Generation with Long Context

2026-02-05 · Shuo Chen, Cong Wei, Sun Sun, Ping Nie, Kai Zhou, Ge Zhang, Ming-Hsuan Yang, Wenhu Chen arxiv

Recent approaches to real-time long video generation typically employ streaming tuning strategies, attempting to train a long-context student using a short-context (memoryless) teacher. In these frameworks, the student performs long rollouts but receives supervision from a teacher limited to short 5-second windows. This structural discrepancy creates a critical \textbf{student-teacher mismatch}: the teacher's inability to access long-term history prevents it from guiding the student on global temporal dependencies, effectively capping the student's context length. To resolve this, we propose \textbf{Context Forcing}, a novel framework that trains a long-context student via a long-context teacher. By ensuring the teacher is aware of the full generation history, we eliminate the supervision mismatch, enabling the robust training of models capable of long-term consistency. To make this computationally feasible for extreme durations (e.g., 2 minutes), we introduce a context management system that transforms the linearly growing context into a \textbf{Slow-Fast Memory} architecture, significantly reducing visual redundancy. Extensive results demonstrate that our method enables effective context lengths exceeding 20 seconds -- 2 to 10 times longer than state-of-the-art methods like LongLive and Infinite-RoPE. By leveraging this extended context, Context Forcing preserves superior consistency across long durations, surpassing state-of-the-art baselines on various long video evaluation metrics.

📄 PDF Abstract BibTeX arXiv:2602.06028

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation

2026-03-22 · Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song 외 arxiv

Autoregressive video diffusion has recently emerged as a promising paradigm for long-video generation, enabling causal synthesis beyond the temporal limits of bidirectional models. Existing forcing-based training strateg…

Video Generation

Pack and Force Your Memory: Long-form and Consistent Video Generation

2025-10-02 · Xiaofei Wu, Guozhen Zhang, Zhiyong Xu, Yuan Zhou 외 arxiv

Long-form video generation presents a dual challenge: models must capture long-range dependencies while preventing the error accumulation inherent in autoregressive decoding. To address these challenges, we make two cont…

Computational EfficiencyVideo Generation

Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity

2026-05-14 · Jiahao Tian, Yiwei Wang, Gang Yu, Chi Zhang arxiv

Autoregressive video diffusion models support real-time synthesis but suffer from error accumulation and context loss over long horizons. We discover that attention heads in AR video diffusion transformers serve function…

Video Generation

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

2025-06-09 · Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou 외

We introduce Self Forcing, a novel training paradigm for autoregressive video diffusion models. It addresses the longstanding issue of exposure bias, where models trained on ground-truth context must generate sequences c…

GPUVideo Generation

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

2026-07-03 · Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat 외 hf

Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models excel at global coherenc…

Video Generation