paper-with-me

Papers

Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation

2021-04-16 · Lincheng Li, Suzhen Wang, Zhimeng Zhang, Yu Ding, Yixing Zheng, Xin Yu, Changjie Fan

In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextual sentiments as well as speech rhythm and pauses. To be specific, our framework consists of a speaker-independent stage and a speaker-specific stage. In the speaker-independent stage, we design three parallel networks to generate animation parameters of the mouth, upper face, and head from texts, separately. In the speaker-specific stage, we present a 3D face model guided attention network to synthesize videos tailored for different individuals. It takes the animation parameters as input and exploits an attention mask to manipulate facial expression changes for the input individuals. Furthermore, to better establish authentic correspondences between visual motions (i.e., facial expression changes and head movements) and audios, we leverage a high-accuracy motion capture dataset instead of relying on long videos of specific individuals. After attaining the visual and audio correspondences, we can effectively train our network in an end-to-end fashion. Extensive experiments on qualitative and quantitative results demonstrate that our algorithm achieves high-quality photo-realistic talking-head videos including various facial expressions and head motions according to speech rhythms and outperforms the state-of-the-art.

📄 PDF Abstract BibTeX arXiv:2104.07995

Code (1)

FuxiVirtualHuman/Write-a-Speaker 공식 구현

Tasks

Face ModelRhythmTalking Head GenerationVideo Generation

Similar Papers 제목 키워드 기반

Three-Stage Speaker Verification Architecture in Emotional Talking Environments

2018-09-03 · Ismail Shahin, Ali Bou Nassif

Speaker verification performance in neutral talking environment is usually high, while it is sharply decreased in emotional talking environments. This performance degradation in emotional environments is due to the probl…

Speaker Verification

CASA-Based Speaker Identification Using Cascaded GMM-CNN Classifier in Noisy and Emotional Talking Conditions

2021-02-11 · Ali Bou Nassif, Ismail Shahin, Shibani Hamsa, Nawel Nemmour 외

This work aims at intensifying text-independent speaker identification performance in real application situations such as noisy and emotional talking conditions. This is achieved by incorporating two different modules: a…

Emotion RecognitionSpeaker Identification

EmoSpeaker: One-shot Fine-grained Emotion-Controlled Talking Face Generation

2024-02-02 · Guanwen Feng, Haoran Cheng, Yunan Li, Zhiyuan Ma 외

Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express v…

AttributeFace GenerationTalking Face Generation

Novel Hybrid DNN Approaches for Speaker Verification in Emotional and Stressful Talking Environments

2021-12-26 · Ismail Shahin, Ali Bou Nassif, Nawel Nemmour, Ashraf Elnagar 외

In this work, we conducted an empirical comparative study of the performance of text-independent speaker verification in emotional and stressful environments. This work combined deep models with shallow architecture, whi…

Speaker VerificationText-Independent Speaker Verification

Talking-head Generation with Rhythmic Head Motion

2020-07-16 · Lele Chen, Guofeng Cui, Celong Liu, Zhong Li 외

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys prosodic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably suc…

Talking Head Generation