Multi-task learning improves synthetic speech detection
With the development of deep learning, synthetic speech has become more and more realistic and easier to spoof Automatic Speaker Verification (ASV) devices. Based on mining more effective hand-crafted features and proposing more powerful networks, many algorithms have been proposed to detect this malicious attack. In this paper, by observing that deepening the network impairs the performance of the network in detecting unknown attacks, we propose that the synthetic speech detection problem is an out-of-distribution (OOD) generalization problem and we enhance the robustness of networks by using multi-task learning. In our system, three auxiliary tasks are used to assist synthetic speech detection: bonafide speech reconstruction, spoofing voice conversion and speaker classification. Experimental results show that our approach can be applied to multiple architectures and can significantly improve the performance on both known attacks (development set) and unknown attacks (evaluation set). In addition, our best-performing network is quite competitive to recent state-of-the-art (SOTA) systems. It demonstrates the potential application of multi-task learning in synthetic speech detection.
Code (1)
Tasks
Multi-Task LearningSpeaker VerificationSynthetic Speech DetectionVoice ConversionSimilar Papers 제목 키워드 기반
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
Current anti-spoofing systems remain vulnerable to expressive and emotional synthetic speech, since they rarely leverage prosody as a discriminative cue. Prosody is central to human expressiveness and emotion, and humans…
Self-Supervised LearningMulti-Task LearningSpoof DetectionHISPASpoof: A New Dataset For Spanish Speech Forensics
Zero-shot Voice Cloning (VC) and Text-to-Speech (TTS) methods have advanced rapidly, enabling the generation of highly realistic synthetic speech and raising serious concerns about their misuse. While numerous detectors …
Multi-Task Adversarial Training Algorithm for Multi-Speaker Neural Text-to-Speech
We propose a novel training algorithm for a multi-speaker neural text-to-speech (TTS) model based on multi-task adversarial training. A conventional generative adversarial network (GAN)-based training algorithm significa…
Generative Adversarial Networktext-to-speechText to SpeechCollaborative Watermarking for Adversarial Speech Synthesis
Advances in neural speech synthesis have brought us technology that is not only close to human naturalness, but is also capable of instant voice cloning with little data, and is highly accessible with pre-trained models …
Speaker VerificationSpeech SynthesisSynthetic Speech DetectionVoice CloningAugmenting Polish Automatic Speech Recognition System With Synthetic Data
This paper presents a system developed for submission to Poleval 2024, Task 3: Polish Automatic Speech Recognition Challenge. We describe Voicebox-based speech synthesis pipeline and utilize it to augment Conformer and W…
Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech Synthesis