paper-with-me

홈 › Papers

DualTalker: A Cross-Modal Dual Learning Approach for Speech-Driven 3D Facial Animation

2023-11-08 · Guinan Su, Yanwu Yang, Zhifeng Li

In recent years, audio-driven 3D facial animation has gained significant attention, particularly in applications such as virtual reality, gaming, and video conferencing. However, accurately modeling the intricate and subtle dynamics of facial expressions remains a challenge. Most existing studies approach the facial animation task as a single regression problem, which often fail to capture the intrinsic inter-modal relationship between speech signals and 3D facial animation and overlook their inherent consistency. Moreover, due to the limited availability of 3D-audio-visual datasets, approaches learning with small-size samples have poor generalizability that decreases the performance. To address these issues, in this study, we propose a cross-modal dual-learning framework, termed DualTalker, aiming at improving data usage efficiency as well as relating cross-modal dependencies. The framework is trained jointly with the primary task (audio-driven facial animation) and its dual task (lip reading) and shares common audio/motion encoder components. Our joint training framework facilitates more efficient data usage by leveraging information from both tasks and explicitly capitalizing on the complementary relationship between facial motion and audio to improve performance. Furthermore, we introduce an auxiliary cross-modal consistency loss to mitigate the potential over-smoothing underlying the cross-modal complementary representations, enhancing the mapping of subtle facial expression dynamics. Through extensive experiments and a perceptual user study conducted on the VOCA and BIWI datasets, we demonstrate that our approach outperforms current state-of-the-art methods both qualitatively and quantitatively. We have made our code and video demonstrations available at https://github.com/sabrina-su/iadf.git.

📄 PDF Abstract BibTeX arXiv:2311.04766

Code (0)

등록된 구현이 없습니다.

Tasks

Lip Reading

Similar Papers 제목 키워드 기반

Uncertainty Modeling in Multimodal Speech Analysis Across the Psychosis Spectrum

2025-02-25 · Morteza Rohanian, Roya M. Hüppi, Farhad Nooralahzadeh, Noemi Dannecker 외

Capturing subtle speech disruptions across the psychosis spectrum is challenging because of the inherent variability in speech patterns. This variability reflects individual differences and the fluctuating nature of symp…

Diagnostic

Dual Audio-Centric Modality Coupling for Talking Head Generation

2025-03-26 · Ao Fu, Ziqi Ni, Yi Zhou

The generation of audio-driven talking head videos is a key challenge in computer vision and graphics, with applications in virtual avatars and digital media. Traditional approaches often struggle with capturing the comp…

NeRFTalking Head Generationtext-to-speechText to Speech

Speech Tokenizer is Key to Consistent Representation

2025-07-09 · Wonjin Jung, Sungil Kang, Dong-Yeon Cho arxiv

Speech tokenization is crucial in digital speech processing, converting continuous speech signals into discrete units for various computational tasks. This paper introduces a novel speech tokenizer with broad applicabili…

Emotion RecognitionVoice Conversion

AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models

2026-05-08 · Kai Zheng, Zejian Kang, Rui Mao, Hongyuan Zou 외 arxiv

Speech-driven facial animation requires accurate correspondence between acoustic signals and facial motion, especially for articulation-related mouth movements. However, directly mapping speech audio to facial coefficien…

Soft Alignment of Modality Space for End-to-end Speech Translation

2023-12-18 · Yuhao Zhang, Kaiqi Kou, Bei Li, Chen Xu 외

End-to-end Speech Translation (ST) aims to convert speech into target text within a unified model. The inherent differences between speech and text modalities often impede effective cross-modal and cross-lingual transfer…

Cross-Lingual TransferTranslation