paper-with-me

Papers

Adapting VACE for Real-Time Autoregressive Video Diffusion

2026-02-16 · Ryan Fosdick arxiv

We describe an adaptation of VACE (Video All-in-one Creation and Editing) for real-time autoregressive video generation. VACE provides unified video control (reference guidance, structural conditioning, inpainting, and temporal extension) but assumes bidirectional attention over full sequences, making it incompatible with streaming pipelines that require fixed chunk sizes and causal attention. The key modification moves reference frames from the diffusion latent space into a parallel conditioning pathway, preserving the fixed chunk sizes and KV caching that autoregressive models require. This adaptation reuses existing pretrained VACE weights without additional training. Across 1.3B and 14B model scales, VACE adds 20-30% latency overhead for structural control and inpainting, with negligible VRAM cost relative to the base model. Reference-to-video fidelity is severely degraded compared to batch VACE due to causal attention constraints. A reference implementation is available at https://github.com/daydreamlive/scope.

📄 PDF Abstract BibTeX arXiv:2602.14381

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VACE: All-in-One Video Creation and Editing

2025-03-10 · Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang 외

Diffusion Transformer has demonstrated powerful capability and scalability in generating high-quality images and videos. Further pursuing the unification of generation and editing tasks has yielded significant progress i…

AllHuman-Domain Subject-to-VideoOpen-Domain Subject-to-VideoSingle-Domain Subject-to-Video+2

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

2026-05-28 · Min Zhao, Hongzhou Zhu, Bokai Yan, Zihan Zhou 외 arxiv

Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains challenging. Interactive world models re…

Video Generation

Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation

2026-05-19 · Matthew Bendel, Stephen W. Bailey, Mithilesh Vaidya, Sumukh Badam 외 arxiv

Long-horizon video generation suffers from two intertwined issues. First, there is drift, where video quality degrades over time. Second, there are continuity issues which manifest as object permanence issues, or imprope…

Video Generation

Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning

2025-10-16 · Xiangyu Meng, Zixian Zhang, Zhenghao Zhang, Junchao Liao 외 arxiv

While advanced methods like VACE and Phantom have advanced video generation for specific subjects in diverse scenarios, they struggle with multi-human identity preservation in dynamic interactions, where consistent ident…

Reinforcement LearningVideo Generation

Proceedings Wivace 2013 - Italian Workshop on Artificial Life and Evolutionary Computation

2013-09-27 · Alex Graudenzi, Giulio Caravagna, Giancarlo Mauri, Marco Antoniotti

The Wivace 2013 Electronic Proceedings in Theoretical Computer Science (EPTCS) contain some selected long and short articles accepted for the presentation at Wivace 2013 - Italian Workshop on Artificial Life and Evolutio…

ArticlesArtificial Life