paper-with-me

홈 › Papers

EchoAvatar: Real-time Generative Avatar Animation from Audio Streams

2026-05-27 · Bohong Chen, Yumeng Li, Yinglin Xu, Youyi Zheng, Yanlin Weng, Kun Zhou arxiv

Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistants. However, most existing approaches are limited to offline processing of complete audio sequences or are constrained to specific domains, rarely handling both speech and music effectively. In this paper, we introduce a novel framework designed to generate continuous, coherent full-body motion from streaming speech and music with low latency. Central to our approach is a unified streaming architecture capable of synthesizing continuous motion from incremental audio inputs. We employ a robust training strategy that enforces strong audio dependency, allowing the model to seamlessly generalize across conversational speech and rhythmic music without requiring explicit domain labels or mode switching. Additionally, we explored Reinforcement Learning to refine the quality of online generation. Furthermore, we bridge reactive animation with intent-driven behavior via a tool-call interface that allows upstream Large Language Models to inject explicit semantic control. By combining this controllability with stream audio-driven synthesis, our framework serves as a plug-and-play solution for transforming voice agents into interactive humanoid avatars. Extensive experiments demonstrate that our method outperforms state-of-the-art realtime baselines in motion quality and synchronization while maintaining the flexibility required for live deployment. Our code, pre-trained models, and videos are available at https://robinwitch.github.io/EchoAvatar-Page.

📄 PDF Abstract BibTeX arXiv:2605.28272

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars

2025-08-22 · NVIDIA, :, Chaeyeon Chung, Ilya Fedorov 외 arxiv

Audio-driven facial animation presents an effective solution for animating digital avatars. In this paper, we detail the technical aspects of NVIDIA Audio2Face-3D, including data acquisition, network architecture, retarg…

One-Shot Feed-Forward 360$^{\circ}$ Animatable Avatar via Inpainted UV-Space Gaussian Modeling

2026-01-19 · Shuling Zhao, Dan Xu arxiv

Building one-shot 3D animatable head avatars is an important yet challenging problem. Existing methods generally collapse under large camera pose variations, compromising the realism of 3D avatars. In this work, we propo…

3D Reconstruction

Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions

2025-04-27 · Mohammad Mahdi Abootorabi, Omid Ghahroodi, Pardis Sadat Zahraei, Hossein Behzadasl 외

Generative AI is reshaping art, gaming, and most notably animation. Recent breakthroughs in foundation and diffusion models have reduced the time and cost of producing animated content. Characters are central animation c…

Image GenerationMotion SynthesisSurveyTexture Synthesis

AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian Reconstruction

2024-12-03 · CVPR 2025 1 · Lingteng Qiu, Shenhao Zhu, Qi Zuo, Xiaodong Gu 외

Generating animatable human avatars from a single image is essential for various digital human modeling applications. Existing 3D reconstruction methods often struggle to capture fine details in animatable models, while …

3D ReconstructionVideo Generation

FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time Animation

2026-01-20 · Xinya Ji, Sebastian Weiss, Manuel Kansy, Jacek Naruniec 외 arxiv

Despite recent progress in 3D Gaussian-based head avatar modeling, efficiently generating high fidelity avatars remains a challenge. Current methods typically rely on extensive multi-view capture setups or monocular vide…