paper-with-me

Papers

Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory

2026-07-29 · Yanbo Ding, Zhizhi Guo, Quanyue Song, Yishan He, Zhixiang He, Yongxiang Li, Yali Wang arxiv

Audio-video generative models achieve impressive quality but suffer from high latency, making them unsuitable for real-time applications. Although several streaming audio-video generation methods have been proposed, they remain costly and fail to support long-form generation. To address this, we propose \textbf{Ripple}, a real-time joint audio-video generation system with a cross-modal recurrent memory mechanism. To enable efficient streaming inference while preserving long-term context, Ripple combines a fixed-length sliding-window attention with modality-specific memory states that continuously summarize audio and video context. Cross-modal memory interaction is further introduced to enhance audio-visual synchronization. To learn this memory-augmented model effectively, we devise a three-stage training recipe: (1) adapting a bidirectional audio-video teacher to block-wise causal attention with simulated memory, (2) optimizing the memory construction and interaction pipeline through end-to-end distillation, and (3) applying online reinforcement post-training tailored for streaming audio-video generation. As a result, Ripple achieves ~28 FPS at 480P resolution, over faster than the teacher, while capable of coherent long-form generation. Extensive experiments on both short-video and long-video benchmarks demonstrate our superior performance over existing offline and online joint audio-video generation methods.

📄 PDF Abstract BibTeX arXiv:2607.26818

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration

2026-05-25 · Linrui Tian, Qi Wang, Bang Zhang arxiv

Real-time streaming joint audio-video generation for character animation requires a generator to speak the requested transcript, maintain visual identity across chunks, and run within a strict playback budget. These requ…

Video GenerationVideo Denoising

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

2026-04-26 · Chunyu Li, Jiaye Li, Ruiqiao Mei, Haoyuan Xia 외 arxiv

Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio-visual diffusion models remain too slow…

Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Video Avatar Generation

2026-08-07 · Lunjie Zhu, Xingtong Ge, Fangyu Lin, Yi Zhang 외 arxiv

Joint audio-video generative models serve as foundation for immersive and interactive digital-human generation. Nevertheless, most existing models rely on bidirectional attention and multi-step denoising and can generate…

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

2026-06-23 · Lianghua Huang, Zhi-Fan Wu, Wei Wang, Yupeng Shi 외 arxiv

We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. Wan-Streamer seamlessly models language, …

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

2026-08-06 · Menglin Han, Yang Ding, Yulei Lu, Haoran Yu 외 arxiv

Real-time long-form avatar audio--video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting pre…

Video Generation