paper-with-me

홈 › Papers

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

2025-01-01 · CVPR 2025 1 · Antoni Bigata, Michał Stypułkowski, Rodrigo Mira, Stella Bounareli, Konstantinos Vougioukas, Zoe Landgraf, Nikita Drobyshev, Maciej Zieba, Stavros Petridis, Maja Pantic

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. Existing methods attempt to mitigate this through external spatial control, increasing long-term consistency but compromising the naturalness of motion. We propose KeyFace, a novel two-stage diffusion-based framework, to address these issues. In the first stage, keyframes are generated at a low frame rate, conditioned on audio input and an identity frame, to capture essential facial expressions and movements over extended periods of time. In the second stage, an interpolation model fills in the gaps between keyframes, ensuring smooth transitions and temporal coherence. To further enhance realism, we incorporate continuous emotion representations and handle a wide range of non-speech vocalizations (NSVs), such as laughter and sighs. We also introduce two new evaluation metrics for assessing lip synchronization and NSV generation. Experimental results show that KeyFace outperforms state-of-the-art methods in generating natural, coherent facial animations over extended durations, successfully encompassing NSVs and continuous emotions.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens

2026-05-29 · Qingcheng Zhao, Yifang Pan, Karan Singh arxiv

Recent advances in Audio-LLMs like GPT-4o have ushered in an era of conversational interaction with language models. Conversational avatars however, still seem robotic in facial expression and conversational flow, in par…

Speech RecognitionSpeech SynthesisText Generation

Joint Audio-Text Model for Expressive Speech-Driven 3D Facial Animation

2021-12-04 · Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang 외

Speech-driven 3D facial animation with accurate lip synchronization has been widely studied. However, synthesizing realistic motions for the entire face during speech has rarely been explored. In this work, we present a …

Language Modelling

3DiFACE: Diffusion-based Speech-driven 3D Facial Animation and Editing

2023-12-01 · Balamurugan Thambiraja, Sadegh Aliakbarian, Darren Cosker, Justus Thies

We present 3DiFACE, a novel method for personalized speech-driven 3D facial animation and editing. While existing methods deterministically predict facial animations from speech, they overlook the inherent one-to-many re…

Diversity

StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model

2025-11-18 · Yifan Yang, Zhi Cen, Sida Peng, Xiangwei Chen 외 arxiv

This paper focuses on the task of speech-driven 3D facial animation, which aims to generate realistic and synchronized facial motions driven by speech inputs. Recent methods have employed audio-conditioned diffusion mode…

X-Actor: Emotional and Expressive Long-Range Portrait Acting from Audio

2025-08-04 · Chenxu Zhang, Zenan Li, Hongyi Xu, You Xie 외 arxiv

We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that e…