paper-with-me

Papers

Diverse Code Query Learning for Speech-Driven Facial Animation

2024-09-27 · Chunzhi Gu, Shigeru Kuriyama, Katsuya Hotta

Speech-driven facial animation aims to synthesize lip-synchronized 3D talking faces following the given speech signal. Prior methods to this task mostly focus on pursuing realism with deterministic systems, yet characterizing the potentially stochastic nature of facial motions has been to date rarely studied. While generative modeling approaches can easily handle the one-to-many mapping by repeatedly drawing samples, ensuring a diverse mode coverage of plausible facial motions on small-scale datasets remains challenging and less explored. In this paper, we propose predicting multiple samples conditioned on the same audio signal and then explicitly encouraging sample diversity to address diverse facial animation synthesis. Our core insight is to guide our model to explore the expressive facial latent space with a diversity-promoting loss such that the desired latent codes for diversification can be ideally identified. To this end, building upon the rich facial prior learned with vector-quantized variational auto-encoding mechanism, our model temporally queries multiple stochastic codes which can be flexibly decoded into a diverse yet plausible set of speech-faithful facial motions. To further allow for control over different facial parts during generation, the proposed model is designed to predict different facial portions of interest in a sequential manner, and compose them to eventually form full-face motions. Our paradigm realizes both diverse and controllable facial animation synthesis in a unified formulation. We experimentally demonstrate that our method yields state-of-the-art performance both quantitatively and qualitatively, especially regarding sample diversity.

📄 PDF Abstract BibTeX arXiv:2409.19143

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior

2023-01-06 · CVPR 2023 1 · Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun 외

Speech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data. Existing works typically formul…

3D Face Animationregression

EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

2023-03-20 · ICCV 2023 1 · Ziqiao Peng, HaoYu Wu, Zhenbo Song, Hao Xu 외

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them fr…

3D Face AnimationDecoderDisentanglement

DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation

2024-08-12 · Jisoo Kim, Jungbin Cho, Joonho Park, Soonmin Hwang 외

Speech-driven 3D facial animation has garnered lots of attention thanks to its broad range of applications. Despite recent advancements in achieving realistic lip motion, current methods fail to capture the nuanced emoti…

3D Face AnimationContrastive Learning

SAiD: Speech-driven Blendshape Facial Animation with Diffusion

2023-12-25 · Inkyu Park, Jaewoong Cho

Speech-driven 3D facial animation is challenging due to the scarcity of large-scale visual-audio datasets despite extensive research. Most prior works, typically focused on learning regression models on a small dataset u…

Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation

2023-12-18 · Hui Fu, Zeqing Wang, Ke Gong, Keze Wang 외

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip s…

DisentanglementRepresentation Learning