paper-with-me

홈 › Papers

An End-to-End Multi-Module Audio Deepfake Generation System for ADD Challenge 2023

2023-07-03 · Sheng Zhao, Qilong Yuan, Yibo Duan, Zhuoyue Chen

The task of synthetic speech generation is to generate language content from a given text, then simulating fake human voice.The key factors that determine the effect of synthetic speech generation mainly include speed of generation, accuracy of word segmentation, naturalness of synthesized speech, etc. This paper builds an end-to-end multi-module synthetic speech generation model, including speaker encoder, synthesizer based on Tacotron2, and vocoder based on WaveRNN. In addition, we perform a lot of comparative experiments on different datasets and various model structures. Finally, we won the first place in the ADD 2023 challenge Track 1.1 with the weighted deception success rate (WDSR) of 44.97%.

📄 PDF Abstract BibTeX arXiv:2307.00729

Code (0)

등록된 구현이 없습니다.

Tasks

Face Swapping

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
WaveRNN WaveRNN is a single-layer recurrent neural network for audio generation that is designed efficiently predict 16-bit raw audio samples. The overall computation in the…

Similar Papers 제목 키워드 기반

Source Tracing of Audio Deepfake Systems

2024-07-10 · Nicholas Klein, Tianxiang Chen, Hemlata Tak, Ricardo Casal 외

Recent progress in generative AI technology has made audio deepfakes remarkably more realistic. While current research on anti-spoofing systems primarily focuses on assessing whether a given audio sample is fake or genui…

Face Swappingtext-to-speechText to SpeechVoice Conversion

Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio

2024-12-02 · Xinrui Yan, Jiangyan Yi, JianHua Tao, Yujie Chen 외

Open environment oriented open set model attribution of deepfake audio is an emerging research topic, aiming to identify the generation models of deepfake audio. Most previous work requires manually setting a rejection t…

Face Swapping

Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework

2025-06-09 · Kuiyuan Zhang, Wenjie Pei, Rushi Lan, Yifang Guo 외

Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly e…

audio-visual learningDeepFake DetectionFace SwappingMisinformation+1

FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset

2021-08-11 · Hasam Khalid, Shahroz Tariq, Minha Kim, Simon S. Woo

While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be us…

DeepFake DetectionFace Swapping

Towards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework

2025-08-04 · Andrea Di Pierno, Luca Guarnera, Dario Allegra, Sebastiano Battiato arxiv

The proliferation of audio deepfakes poses a growing threat to trust in digital communications. While detection methods have advanced, attributing audio deepfakes to their source models remains an underexplored yet cruci…

Audio Deepfake Detection