paper-with-me

Papers

Model See Model Do: Speech-Driven Facial Animation with Style Control

2025-05-02 · Yifang Pan, Karan Singh, Luiz Gustavo Hafemann

Speech-driven 3D facial animation plays a key role in applications such as virtual avatars, gaming, and digital content creation. While existing methods have made significant progress in achieving accurate lip synchronization and generating basic emotional expressions, they often struggle to capture and effectively transfer nuanced performance styles. We propose a novel example-based generation framework that conditions a latent diffusion model on a reference style clip to produce highly expressive and temporally coherent facial animations. To address the challenge of accurately adhering to the style reference, we introduce a novel conditioning mechanism called style basis, which extracts key poses from the reference and additively guides the diffusion generation process to fit the style without compromising lip synchronization quality. This approach enables the model to capture subtle stylistic cues while ensuring that the generated animations align closely with the input speech. Extensive qualitative, quantitative, and perceptual evaluations demonstrate the effectiveness of our method in faithfully reproducing the desired style while achieving superior lip synchronization across various speech scenarios.

📄 PDF Abstract BibTeX arXiv:2505.01319

Code (0)

등록된 구현이 없습니다.

Tasks

model

Methods 이 논문이 사용한 방법론

Latent Diffusion Model Diffusion models applied to latent spaces, which are normally built with (Variational) Autoencoders.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Personalized Speech-driven Expressive 3D Facial Animation Synthesis with Style Control

2023-10-25 · Elif Bozkurt

Different people have different facial expressions while speaking emotionally. A realistic facial animation system should consider such identity-specific speaking styles and facial idiosyncrasies to achieve high-degree o…

Decoder

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization

2026-06-26 · Arthur Josi, Emeline Got, Abdallah Dib, Luiz Gustavo Hafemann 외 arxiv

Speech-driven 3D facial animation methods face significant challenges in simultaneously achieving high-fidelity motion and precise artistic control at production quality. Existing controllable models typically learn glob…

ExpCLIP: Bridging Text and Facial Expressions via Semantic Alignment

2023-08-28 · Yicheng Zhong, Huawei Wei, Peiji Yang, Zhisheng Wang

The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression tem…

Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation

2023-12-18 · Hui Fu, Zeqing Wang, Ke Gong, Keze Wang 외

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip s…

DisentanglementRepresentation Learning

CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation

2026-05-28 · Xuangeng Chu, Yuan Gan, Ziteng Cui, Shuhong Liu 외 arxiv

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predef…

Talking Head Generation