paper-with-me

홈 › Papers

Attention-Based Lip Audio-Visual Synthesis for Talking Face Generation in the Wild

2022-03-08 · Ganglai Wang, Peng Zhang, Lei Xie, Wei Huang, Yufei zha

Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standing challenge to be further investigated. Motivated by xxx, in this paper, an AttnWav2Lip model is proposed by incorporating spatial attention module and channel attention module into lip-syncing strategy. Rather than focusing on the unimportant regions of the face image, the proposed AttnWav2Lip model is able to pay more attention on the lip region reconstruction. To our limited knowledge, this is the first attempt to introduce attention mechanism to the scheme of talking face generation. An extensive experiments have been conducted to evaluate the effectiveness of the proposed model. Compared to the baseline measured by LSE-D and LSE-C metrics, a superior performance has been demonstrated on the benchmark lip synthesis datasets, including LRW, LRS2 and LRS3.

📄 PDF Abstract BibTeX arXiv:2203.03984

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationTalking Face Generation

Methods 이 논문이 사용한 방법론

Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Average Pooling 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Sigmoid Activation 설명 없음

Similar Papers 제목 키워드 기반

Neural Text to Articulate Talk: Deep Text to Audiovisual Speech Synthesis achieving both Auditory and Photo-realism

2023-12-11 · Georgios Milis, Panagiotis P. Filntisis, Anastasios Roussos, Petros Maragos

Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-sync…

Face GenerationLip ReadingSpeech SynthesisTalking Face Generation+2

Text-driven Talking Face Synthesis by Reprogramming Audio-driven Models

2023-06-28 · Jeongsoo Choi, Minsu Kim, Se Jin Park, Yong Man Ro

In this paper, we present a method for reprogramming pre-trained audio-driven talking face synthesis models to operate in a text-driven manner. Consequently, we can easily generate face videos that articulate the provide…

Face Generation

NeRF-AD: Neural Radiance Field with Attention-based Disentanglement for Talking Face Synthesis

2024-01-23 · Chongke Bi, Xiaoxing Liu, Zhilei Liu

Talking face synthesis driven by audio is one of the current research hotspots in the fields of multidimensional signal processing and multimedia. Neural Radiance Field (NeRF) has recently been brought to this research f…

DisentanglementFace GenerationNeRF

Arbitrary Talking Face Generation via Attentional Audio-Visual Coherence Learning

2018-12-17 · Hao Zhu, Huaibo Huang, Yi Li, Aihua Zheng 외

Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing met…

Face GenerationTalking Face Generation

JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing

2025-01-03 · Qili Wang, Dajiang Wu, Zihang Xu, Junshi Huang 외

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper i…

3D ReconstructionFace GenerationMotion GenerationTalking Face Generation+2