paper-with-me

Papers

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

2026-05-27 · Kaiyang Ji, Bingsheng Qian, Binghuan Wu, Kangyi Chen, Ye Shi, Jingya Wang arxiv

We study real-time audio-responsive character control as a deployment-faithful problem: strictly causal, bounded-latency streaming that must generate coherent full-body motion at interactive frame rates while the audio condition can change abruptly, including tempo shifts, drops, or user edits. Prior music-to-motion systems are largely optimized for offline generation with global context, and degrade in streaming rollouts where conditioning history becomes stale or unreliable. We introduce DiscoForcing, a streaming audio-driven diffusion framework that combines a causal music encoder that captures rhythmic structure and phase dynamics with a diffusion-forcing sequence model trained under heterogeneous noise levels across the temporal horizon. Building on this, we design a hybrid temporal schedule and a history-guided streaming sampler to explicitly trade off responsiveness against long-horizon consistency under non-stationary audio. Implemented in an end-to-end real-time interactive system with online avatar playback and humanoid deployment workflows, DiscoForcing delivers more stable long-horizon rollouts and sharper audio-motion alignment than prior baselines under matched causality and latency constraints while maintaining real-time throughput.

📄 PDF Abstract BibTeX arXiv:2605.28491

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

2026-07-25 · Quanyue Song, Yishan He, Yanbo Ding, Zhixiang He 외 arxiv

Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for interactive avatar systems. However, ext…

UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models

2026-01-04 · Qundong Shi, Jie Zhou, Biyuan Lin, Junbo Cui 외 arxiv

The development of audio foundation models has accelerated rapidly since the emergence of GPT-4o. However, the lack of comprehensive evaluation has become a critical bottleneck for further progress in the field, particul…

Audio Generation

UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars

2026-05-14 · Xiaoyu Zhan, Xinyu Fu, Chenghao Yang, Xiaohong Zhang 외 arxiv

Speech-driven gestures and facial animations are fundamental to expressive digital avatars in games, virtual production, and interactive media. However, existing methods are either limited to a single modality for audio …

Unified AI for Accurate Audio Anomaly Detection

2025-05-20 · Hamideh Khaleghpour, Brett McKinney

This paper presents a unified AI framework for high-accuracy audio anomaly detection by integrating advanced noise reduction, feature extraction, and machine learning modeling techniques. The approach combines spectral s…

Anomaly Detection

RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer

2025-08-07 · Fangyu Du, Taiqing Li, Qian Qiao, Tan Yu 외 arxiv

Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high…