paper-with-me

Papers

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

2025-06-03 · Chetwin Low, Weimin WANG

In this paper, we present TalkingMachines -- an efficient framework that transforms pretrained video generation models into real-time, audio-driven character animators. TalkingMachines enables natural conversational experiences by integrating an audio large language model (LLM) with our video generation foundation model. Our primary contributions include: (1) We adapt a pretrained SOTA image-to-video DiT into an audio-driven avatar generation model of 18 billion parameters; (2) We enable infinite video streaming without error accumulation through asymmetric knowledge distillation from a bidirectional teacher model into a sparse causal, autoregressive student model; (3) We design a high-throughput, low-latency inference pipeline incorporating several key engineering optimizations such as: (a) disaggregation of the DiT and VAE decoder across separate devices, (b) efficient overlap of inter-device communication and computation using CUDA streams, (c) elimination of redundant recomputations to maximize frame-generation throughput. Please see demo videos here - https://aaxwaz.github.io/TalkingMachines/

📄 PDF Abstract BibTeX arXiv:2506.03099

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderKnowledge DistillationLanguage ModelingLanguage ModellingLarge Language ModelVideo Generation

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

High-fidelity Face Tracking for AR/VR via Deep Lighting Adaptation

2021-03-29 · CVPR 2021 1 · Lele Chen, Chen Cao, Fernando de la Torre, Jason Saragih 외

3D video avatars can empower virtual communications by providing compression, privacy, entertainment, and a sense of presence in AR/VR. Best 3D photo-realistic AR/VR avatars driven by video, that can minimize uncanny eff…

Vocal Bursts Intensity Prediction

RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer

2025-08-07 · Fangyu Du, Taiqing Li, Qian Qiao, Tan Yu 외 arxiv

Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high…

SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation

2025-01-24 · Yujian Liu, Shidang Xu, Jing Guo, Dingbin Wang 외

Generating talking avatar driven by audio remains a significant challenge. Existing methods typically require high computational costs and often lack sufficient facial detail and realism, making them unsuitable for appli…

NeRF

Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars

2025-08-22 · NVIDIA, :, Chaeyeon Chung, Ilya Fedorov 외 arxiv

Audio-driven facial animation presents an effective solution for animating digital avatars. In this paper, we detail the technical aspects of NVIDIA Audio2Face-3D, including data acquisition, network architecture, retarg…

EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation

2025-10-03 · Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun 외 arxiv

This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3-5 minutes of …

Talking Head Generation