paper-with-me

Papers

Imitating Arbitrary Talking Style for Realistic Audio-DrivenTalking Face Synthesis

2021-10-30 · Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou, Chao Duan, Qingshan Deng

People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the "solemn" style is more standardized and seldomly exhibits exaggerated motions. Due to such huge differences between different styles, it is necessary to incorporate the talking style into audio-driven talking face synthesis framework. In this paper, we propose to inject style into the talking face synthesis framework through imitating arbitrary talking style of the particular reference video. Specifically, we systematically investigate talking styles with our collected \textit{Ted-HD} dataset and construct style codes as several statistics of 3D morphable model~(3DMM) parameters. Afterwards, we devise a latent-style-fusion~(LSF) model to synthesize stylized talking faces by imitating talking styles from the style codes. We emphasize the following novel characteristics of our framework: (1) It doesn't require any annotation of the style, the talking style is learned in an unsupervised manner from talking videos in the wild. (2) It can imitate arbitrary styles from arbitrary videos, and the style codes can also be interpolated to generate new styles. Extensive experiments demonstrate that the proposed framework has the ability to synthesize more natural and expressive talking styles compared with baseline methods.

📄 PDF Abstract BibTeX arXiv:2111.00203

Code (1)

wuhaozhe/style_avatar 공식 구현 pytorch

Tasks

Face Generation

Similar Papers 제목 키워드 기반

Style-Preserving Lip Sync via Audio-Aware Style Reference

2024-08-10 · Weizhi Zhong, Jichang Li, Yinqi Cai, Ming Li 외

Audio-driven lip sync has recently drawn significant attention due to its widespread application in the multimedia domain. Individuals exhibit distinct lip shapes when speaking the same utterance, attributed to the uniqu…

Style Transfer for 2D Talking Head Animation

2023-03-17 · Trong-Thang Pham, Nhat Le, Tuong Do, Hung Nguyen 외

Audio-driven talking head animation is a challenging research topic with many real-world applications. Recent works have focused on creating photo-realistic 2D animation, while learning different talking or singing style…

Style Transfer

StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation

2022-08-23 · Dongchan Min, Minyoung Song, Eunji Ko, Sung Ju Hwang

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and …

Talking Head GenerationVideo Generation

High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space Learning

2023-05-04 · CVPR 2023 1 · Chao Xu, Junwei Zhu, Jiangning Zhang, Yue Han 외

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical appl…

Face GenerationTalking Face Generation

DREAM-Talk: Diffusion-based Realistic Emotional Audio-driven Method for Single Image Talking Face Generation

2023-12-21 · Chenxu Zhang, Chao Wang, Jianfeng Zhang, Hongyi Xu 외

The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as exp…

Face GenerationTalking Face Generation