paper-with-me

홈 › Papers

Streaming Generation of Co-Speech Gestures via Accelerated Rolling Diffusion

2025-03-13 · Evgeniia Vu, Andrei Boiarov, Dmitry Vetrov

Generating co-speech gestures in real time requires both temporal coherence and efficient sampling. We introduce Accelerated Rolling Diffusion, a novel framework for streaming gesture generation that extends rolling diffusion models with structured progressive noise scheduling, enabling seamless long-sequence motion synthesis while preserving realism and diversity. We further propose Rolling Diffusion Ladder Acceleration (RDLA), a new approach that restructures the noise schedule into a stepwise ladder, allowing multiple frames to be denoised simultaneously. This significantly improves sampling efficiency while maintaining motion consistency, achieving up to a 2x speedup with high visual fidelity and temporal coherence. We evaluate our approach on ZEGGS and BEAT, strong benchmarks for real-world applicability. Our framework is universally applicable to any diffusion-based gesture generation model, transforming it into a streaming approach. Applied to three state-of-the-art methods, it consistently outperforms them, demonstrating its effectiveness as a generalizable and efficient solution for real-time, high-fidelity co-speech gesture synthesis.

📄 PDF Abstract BibTeX arXiv:2503.10488

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityGesture GenerationMotion SynthesisScheduling

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

2026-07-22 · Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen 외 hf

Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are r…

DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures

2024-09-11 · Steven Hogue, Chenxu Zhang, Hamza Daruger, Yapeng Tian 외

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate takin…

DiversityTalking Head GenerationVideo Generation

The Dynamic Articulatory Model DYNARTmo: Dynamic Movement Generation and Speech Gestures

2025-11-11 · Bernd J. Kröger arxiv

This paper describes the current implementation of the dynamic articulatory model DYNARTmo, which generates continuous articulator movements based on the concept of speech gestures and a corresponding gesture score. The …

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars

2026-06-29 · Habin Lim, Jae-Ho Lee, Hah Min Lew, Ji-Su Kang 외 arxiv

Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. Existing systems only partially address this problem: speech-only full-duplex models can generate speech in…

LivelySpeaker: Towards Semantic-Aware Co-Speech Gesture Generation

2023-09-17 · ICCV 2023 1 · YiHao Zhi, Xiaodong Cun, Xuelin Chen, Xi Shen 외

Gestures are non-verbal but important behaviors accompanying people's speech. While previous methods are able to generate speech rhythm-synchronized gestures, the semantic context of the speech is generally lacking in th…

Gesture GenerationRhythm