paper-with-me

Papers

SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent Space

2024-05-09 · Zeren Zhang, Haibo Qin, Jiayu Huang, Yixin Li, Hui Lin, Yitao Duan, Jinwen Ma

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between tasks and reduce video clarity because the interaction space is limited to the low-level semantic RGB space. To address this issue, we propose an innovative unified framework, SwapTalk, which accomplishes both face swapping and lip synchronization tasks in the same latent space. Referring to recent work on face generation, we choose the VQ-embedding space due to its excellent editability and fidelity performance. To enhance the framework's generalization capabilities for unseen identities, we incorporate identity loss during the training of the face swapping module. Additionally, we introduce expert discriminator supervision within the latent space during the training of the lip synchronization module to elevate synchronization quality. In the evaluation phase, previous studies primarily focused on the self-reconstruction of lip movements in synchronous audio-visual videos. To better approximate real-world applications, we expand the evaluation scope to asynchronous audio-video scenarios. Furthermore, we introduce a novel identity consistency metric to more comprehensively assess the identity consistency over time series in generated facial videos. Experimental results on the HDTF demonstrate that our method significantly surpasses existing techniques in video quality, lip synchronization accuracy, face swapping fidelity, and identity consistency. Our demo is available at http://swaptalk.cc.

📄 PDF Abstract BibTeX arXiv:2405.05636

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationFace SwappingTalking Face Generation

Similar Papers 제목 키워드 기반

JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing

2025-01-03 · Qili Wang, Dajiang Wu, Zihang Xu, Junshi Huang 외

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper i…

3D ReconstructionFace GenerationMotion GenerationTalking Face Generation+2

Audio-Driven Talking Face Generation with Diverse yet Realistic Facial Animations

2023-04-18 · Rongliang Wu, Yingchen Yu, Fangneng Zhan, Jiahui Zhang 외

Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to th…

Face GenerationTalking Face Generation

Audio-Driven Talking Face Video Generation with Dynamic Convolution Kernels

2022-01-16 · Zipeng Ye, Mengfei Xia, Ran Yi, Juyong Zhang 외

In this paper, we present a dynamic convolution kernel (DCK) strategy for convolutional neural networks. Using a fully convolutional network with the proposed DCKs, high-quality talking-face video can be generated from m…

Video Generation

KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation

2024-09-09 · Hoang-Son Vo-Thanh, Quang-Vinh Nguyen, Soo-Hyung Kim

Audio-driven talking face generation is a widely researched topic due to its high applicability. Reconstructing a talking face using audio significantly contributes to fields such as education, healthcare, online convers…

Face GenerationSpeech to Facial LandmarkTalking Face Generation

Speech Driven Talking Face Generation from a Single Image and an Emotion Condition

2020-08-08 · Sefik Emre Eskimez, You Zhang, Zhiyao Duan

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifical…

Emotion RecognitionFace GenerationTalking Face Generation