An End-to-End Multi-Module Audio Deepfake Generation System for ADD Challenge 2023
The task of synthetic speech generation is to generate language content from a given text, then simulating fake human voice.The key factors that determine the effect of synthetic speech generation mainly include speed of generation, accuracy of word segmentation, naturalness of synthesized speech, etc. This paper builds an end-to-end multi-module synthetic speech generation model, including speaker encoder, synthesizer based on Tacotron2, and vocoder based on WaveRNN. In addition, we perform a lot of comparative experiments on different datasets and various model structures. Finally, we won the first place in the ADD 2023 challenge Track 1.1 with the weighted deception success rate (WDSR) of 44.97%.
Code (0)
등록된 구현이 없습니다.
Tasks
Face SwappingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Source Tracing of Audio Deepfake Systems
Recent progress in generative AI technology has made audio deepfakes remarkably more realistic. While current research on anti-spoofing systems primarily focuses on assessing whether a given audio sample is fake or genui…
Face Swappingtext-to-speechText to SpeechVoice ConversionReject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio
Open environment oriented open set model attribution of deepfake audio is an emerging research topic, aiming to identify the generation models of deepfake audio. Most previous work requires manually setting a rejection t…
Face SwappingLightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly e…
audio-visual learningDeepFake DetectionFace SwappingMisinformation+1FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset
While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be us…
DeepFake DetectionFace SwappingTowards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework
The proliferation of audio deepfakes poses a growing threat to trust in digital communications. While detection methods have advanced, attributing audio deepfakes to their source models remains an underexplored yet cruci…
Audio Deepfake Detection