paper-with-me

홈 › Papers

End-to-End Speech-Driven Facial Animation with Temporal GANs

2018-05-23 · Konstantinos Vougioukas, Stavros Petridis, Maja Pantic

Speech-driven facial animation is the process which uses speech signals to automatically synthesize a talking character. The majority of work in this domain creates a mapping from audio features to visual features. This often requires post-processing using computer graphics techniques to produce realistic albeit subject dependent results. We present a system for generating videos of a talking head, using a still image of a person and an audio clip containing speech, that doesn't rely on any handcrafted intermediate features. To the best of our knowledge, this is the first method capable of generating subject independent realistic videos directly from raw audio. Our method can generate videos which have (a) lip movements that are in sync with the audio and (b) natural facial expressions such as blinks and eyebrow movements. We achieve this by using a temporal GAN with 2 discriminators, which are capable of capturing different aspects of the video. The effect of each component in our system is quantified through an ablation study. The generated videos are evaluated based on their sharpness, reconstruction quality, and lip-reading accuracy. Finally, a user study is conducted, confirming that temporal GANs lead to more natural sequences than a static GAN-based approach.

📄 PDF Abstract BibTeX arXiv:1805.09313

Code (1)

PrashanthaTP/wav2mov pytorch

Tasks

Lip Reading

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Realistic Speech-Driven Facial Animation with GANs

2019-06-14 · Konstantinos Vougioukas, Stavros Petridis, Maja Pantic

Speech-driven facial animation is the process that automatically synthesizes talking characters based on speech signals. The majority of work in this domain creates a mapping from audio features to visual features. This …

Audio-Visual SynchronizationLip Reading

Speech-Driven 3D Face Animation with Composite and Regional Facial Movements

2023-08-10 · Haozhe Wu, Songtao Zhou, Jia Jia, Junliang Xing 외

Speech-driven 3D face animation poses significant challenges due to the intricacy and variability inherent in human facial movements. This paper emphasizes the importance of considering both the composite and regional na…

3D Face Animation

Speech-driven Facial Animation using Cascaded GANs for Learning of Motion and Texture

2020-08-01 · ECCV 2020 8 · Dipanjan Das, Sandika Biswas, Sanjana Sinha, Brojeshwar Bhowmick

Speech-driven facial animation methods should produce accurate and realistic lip motions with natural expressions and realistic texture portraying target-specific facial characteristics. Moreover, the methods should also…

Meta-Learning

SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion Diarization

2026-04-14 · Farzaneh Jafari, Stefano Berretti, Anup Basu arxiv

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely…

Talking Head GenerationEmotion Recognition

Personalized Speech-driven Expressive 3D Facial Animation Synthesis with Style Control

2023-10-25 · Elif Bozkurt

Different people have different facial expressions while speaking emotionally. A realistic facial animation system should consider such identity-specific speaking styles and facial idiosyncrasies to achieve high-degree o…

Decoder