paper-with-me

홈 › Papers

SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory

2022-11-02 · Se Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi, Yong Man Ro

The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation learning or leverage intermediate structural information such as landmarks and 3D models. However, they struggle to synthesize fine details of the lips varying at the phoneme level as they do not sufficiently provide visual information of the lips at the video synthesis step. To overcome this limitation, our work proposes Audio-Lip Memory that brings in visual information of the mouth region corresponding to input audio and enforces fine-grained audio-visual coherence. It stores lip motion features from sequential ground truth images in the value memory and aligns them with corresponding audio features so that they can be retrieved using audio input at inference time. Therefore, using the retrieved lip motion features as visual hints, it can easily correlate audio with visual dynamics in the synthesis step. By analyzing the memory, we demonstrate that unique lip features are stored in each memory slot at the phoneme level, capturing subtle lip motion based on memory addressing. In addition, we introduce visual-visual synchronization loss which can enhance lip-syncing performance when used along with audio-visual synchronization loss in our model. Extensive experiments are performed to verify that our method generates high-quality video with mouth shapes that best align with the input audio, outperforming previous state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2211.00924

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual SynchronizationFace GenerationRepresentation LearningTalking Face Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild

2022-11-27 · Kun Cheng, Xiaodong Cun, Yong Zhang, Menghan Xia 외

We present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disen…

Video EditingVideo Generation

Taiwanese-Accented Mandarin and English Multi-Speaker Talking-Face Synthesis System

2022-11-01 · ROCLING 2022 11 · Chia-Hsuan Lin, Jian-Peng Liao, Cho-Chun Hsieh, Kai-Chun Liao 외

This paper proposes a multi-speaker talking-face synthesis system. The system incorporates voice cloning and lip-syncing technology to achieve text-to-talking-face generation by acquiring audio and video clips of any spe…

Face GenerationSpeech SynthesisTalking Face GenerationTransfer Learning+1

Intelligent Video Editing: Incorporating Modern Talking Face Generation Algorithms in a Video Editor

2021-10-16 · Anchit Gupta, Faizan Farooq Khan, Rudrabha Mukhopadhyay, Vinay P. Namboodiri 외

This paper proposes a video editor based on OpenShot with several state-of-the-art facial video editing algorithms as added functionalities. Our editor provides an easy-to-use interface to apply modern lip-syncing algori…

Face GenerationTalking Face GenerationVideo EditingVideo Generation

Attention-Based Lip Audio-Visual Synthesis for Talking Face Generation in the Wild

2022-03-08 · Ganglai Wang, Peng Zhang, Lei Xie, Wei Huang 외

Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standing challenge to be further investigated. …

Face GenerationTalking Face Generation

Neural Text to Articulate Talk: Deep Text to Audiovisual Speech Synthesis achieving both Auditory and Photo-realism

2023-12-11 · Georgios Milis, Panagiotis P. Filntisis, Anastasios Roussos, Petros Maragos

Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-sync…

Face GenerationLip ReadingSpeech SynthesisTalking Face Generation+2