paper-with-me

Papers

Audio-Driven Dubbing for User Generated Contents via Style-Aware Semi-Parametric Synthesis

2023-08-31 · Linsen Song, Wayne Wu, Chaoyou Fu, Chen Change Loy, Ran He

Existing automated dubbing methods are usually designed for Professionally Generated Content (PGC) production, which requires massive training data and training time to learn a person-specific audio-video mapping. In this paper, we investigate an audio-driven dubbing method that is more feasible for User Generated Content (UGC) production. There are two unique challenges to design a method for UGC: 1) the appearances of speakers are diverse and arbitrary as the method needs to generalize across users; 2) the available video data of one speaker are very limited. In order to tackle the above challenges, we first introduce a new Style Translation Network to integrate the speaking style of the target and the speaking content of the source via a cross-modal AdaIN module. It enables our model to quickly adapt to a new speaker. Then, we further develop a semi-parametric video renderer, which takes full advantage of the limited training data of the unseen speaker via a video-level retrieve-warp-refine pipeline. Finally, we propose a temporal regularization for the semi-parametric renderer, generating more continuous videos. Extensive experiments show that our method generates videos that accurately preserve various speaking styles, yet with considerably lower amount of training data and training time in comparison to existing methods. Besides, our method achieves a faster testing speed than most recent methods.

📄 PDF Abstract BibTeX arXiv:2309.00030

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation

2024-06-15 · Ruibo Fu, Shuchen Shi, Hongming Guo, Tao Wang 외

Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape. Despite advancements in AIGC technologies for text and image …

AudioCapsImage Generation

Identity-Preserving Video Dubbing Using Motion Warping

2025-01-08 · Runzhen Liu, Qinjie Lin, Yunfei Liu, Lijian Lin 외

Video dubbing aims to synthesize realistic, lip-synced videos from a reference video and a driving audio signal. Although existing methods can accurately generate mouth shapes driven by audio, they often fail to preserve…

Neural Voice Puppetry: Audio-driven Facial Reenactment

2019-12-11 · ECCV 2020 8 · Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt 외

We present Neural Voice Puppetry, a novel approach for audio-driven facial video synthesis. Given an audio sequence of a source person or digital assistant, we generate a photo-realistic output video of a target person t…

Face ModelNeural RenderingTalking Face GenerationTalking Head Generation+2

Cross-lingual Prosody Transfer for Expressive Machine Dubbing

2023-06-20 · Jakub Swiatkowski, Duo Wang, Mikolaj Babianski, Patrick Lumban Tobing 외

Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to l…

Expressive Speech SynthesisSpeech Synthesis

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

2025-08-19 · Shaoshu Yang, Zhe Kong, Feng Gao, Meng Cheng 외 arxiv

Recent breakthroughs in video AIGC have ushered in a transformative era for audio-driven human animation. However, conventional video dubbing techniques remain constrained to mouth region editing, resulting in discordant…

Video Generation