paper-with-me

홈 › Papers

Audio Deep Fake Detection System with Neural Stitching for ADD 2022

2022-04-19 · Rui Yan, Cheng Wen, Shuran Zhou, Tingwei Guo, Wei Zou, Xiangang Li

This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge\cite{Yi2022ADD}. The very same system was used for both two rounds of evaluation in Track 3.2 with a similar training methodology. The first round of Track 3.2 data is generated from Text-to-Speech(TTS) or voice conversion (VC) algorithms, while the second round of data consists of generated fake audio from other participants in Track 3.1, aiming to spoof our systems. Our systems use a standard 34-layer ResNet, with multi-head attention pooling \cite{india2019self} to learn the discriminative embedding for fake audio and spoof detection. We further utilize neural stitching to boost the model's generalization capability in order to perform equally well in different tasks, and more details will be explained in the following sessions. The experiments show that our proposed method outperforms all other systems with a 10.1% equal error rate(EER) in Track 3.2.

📄 PDF Abstract BibTeX arXiv:2204.08720

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to SpeechVoice Conversion

Methods 이 논문이 사용한 방법론

Attention Pooling 설명 없음
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Average Pooling 설명 없음
Batch Normalization 설명 없음
Kaiming Initialization 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

Can deepfakes be created by novice users?

2023-04-28 · Pulak Mehta, Gauri Jagatap, Kevin Gallagher, Brian Timmerman 외

Recent advancements in machine learning and computer vision have led to the proliferation of Deepfakes. As technology democratizes over time, there is an increasing fear that novice users can create Deepfakes, to discred…

DeepFake DetectionFace Swapping

ADD 2022: the First Audio Deep Synthesis Detection Challenge

2022-02-17 · Jiangyan Yi, Ruibo Fu, JianHua Tao, Shuai Nie 외

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection…

Audio Deepfake DetectionAudio GenerationDeepFake DetectionFace Swapping

Towards the Development of a Real-Time Deepfake Audio Detection System in Communication Platforms

2024-03-18 · Jonat John Mathew, Rakin Ahsan, Sae Furukawa, Jagdish Gautham Krishna Kumar 외

Deepfake audio poses a rising threat in communication platforms, necessitating real-time detection for audio stream integrity. Unlike traditional non-real-time approaches, this study assesses the viability of employing s…

Face Swapping

EnvSDD: Benchmarking Environmental Sound Deepfake Detection

2025-05-25 · Han Yin, Yang Xiao, Rohan Kumar Das, Jisheng Bai 외

Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech or singing voice. However, environmental …

Audio Deepfake DetectionAudio GenerationBenchmarkingDeepFake Detection+1

Multilingual Dataset Integration Strategies for Robust Audio Deepfake Detection: A SAFE Challenge System

2025-08-28 · Hashim Ali, Surya Subramani, Lekha Bollinani, Nithin Sai Adupa 외 arxiv

The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-su…

Self-Supervised LearningAudio Deepfake Detection