paper-with-me

Papers

StableFace: Analyzing and Improving Motion Stability for Talking Face Generation

2022-08-29 · Jun Ling, Xu Tan, Liyang Chen, Runnan Li, Yuchao Zhang, Sheng Zhao, Li Song

While previous speech-driven talking face generation methods have made significant progress in improving the visual quality and lip-sync quality of the synthesized videos, they pay less attention to lip motion jitters which greatly undermine the realness of talking face videos. What causes motion jitters, and how to mitigate the problem? In this paper, we conduct systematic analyses on the motion jittering problem based on a state-of-the-art pipeline that uses 3D face representations to bridge the input audio and output video, and improve the motion stability with a series of effective designs. We find that several issues can lead to jitters in synthesized talking face video: 1) jitters from the input 3D face representations; 2) training-inference mismatch; 3) lack of dependency modeling among video frames. Accordingly, we propose three effective solutions to address this issue: 1) we propose a gaussian-based adaptive smoothing module to smooth the 3D face representations to eliminate jitters in the input; 2) we add augmented erosions on the input data of the neural renderer in training to simulate the distortion in inference to reduce mismatch; 3) we develop an audio-fused transformer generator to model dependency among video frames. Besides, considering there is no off-the-shelf metric for measuring motion jitters in talking face video, we devise an objective metric (Motion Stability Index, MSI), to quantitatively measure the motion jitters by calculating the reciprocal of variance acceleration. Extensive experimental results show the superiority of our method on motion-stable face video generation, with better quality than previous systems.

📄 PDF Abstract BibTeX arXiv:2208.13717

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationTalking Face GenerationVideo Generation

Similar Papers 제목 키워드 기반

EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot Personalization

2026-03-22 · Haolan Xu, Keli Cheng, Lei Wang, Ning Bi 외 arxiv

Audio-driven 3D talking head synthesis has advanced rapidly with Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). By leveraging rich pre-trained priors, few-shot methods enable instant personalization from…

EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion

2024-11-23 · CVPR 2025 1 · Haotian Wang, Yuzhe Weng, Yueyan Li, Zilu Guo 외

Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk …

Talking Head Generation

EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model

2022-05-30 · Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu 외

Although significant progress has been made to audio-driven talking face generation, existing methods either neglect facial emotion or cannot be applied to arbitrary subjects. In this paper, we propose the Emotion-Aware …

Face GenerationTalking Face Generation

High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model

2024-08-10 · Weizhi Zhong, Junfan Lin, Peixin Chen, Liang Lin 외

Audio-driven talking face video generation has attracted increasing attention due to its huge industrial potential. Some previous methods focus on learning a direct mapping from audio to visual content. Despite progress,…

Face GenerationTalking Face GenerationVideo Generation

MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation

2020-08-01 · ECCV 2020 8 · Kaisiyuan Wang Qianyi Wu Linsen Song Zhuoqian Yang Wayne Wu Chen Qian Ran He Yu Qiao Chen Change Loy

The synthesis of natural emotional reactions is an essentialcriteria in vivid talking-face video generation. This criteria is nevertheless seldom taken into consideration in previous works due to the absence of a large-s…

Face GenerationTalking Face GenerationTalking Head GenerationVideo Generation