A Practical Guide to Logical Access Voice Presentation Attack Detection
Voice-based human-machine interfaces with an automatic speaker verification (ASV) component are commonly used in the market. However, the threat from presentation attacks is also growing since attackers can use recent speech synthesis technology to produce a natural-sounding voice of a victim. Presentation attack detection (PAD) for ASV, or speech anti-spoofing, is therefore indispensable. Research on voice PAD has seen significant progress since the early 2010s, including the advancement in PAD models, benchmark datasets, and evaluation campaigns. This chapter presents a practical guide to the field of voice PAD, with a focus on logical access attacks using text-to-speech and voice conversion algorithms and spoofing countermeasures based on artifact detection. It introduces the basic concept of voice PAD, explains the common techniques, and provides an experimental study using recent methods on a benchmark dataset. Code for the experiments is open-sourced.
Code (1)
Tasks
Artifact DetectionSpeaker VerificationSpeech Synthesistext-to-speechText to SpeechVoice ConversionSimilar Papers 제목 키워드 기반
One-class learning towards generalized voice spoofing detection
Human voices can be used to authenticate the identity of the speaker, but the automatic speaker verification (ASV) systems are vulnerable to voice spoofing attacks, such as impersonation, replay, text-to-speech, and voic…
Speaker Verificationtext-to-speechText to SpeechVoice Anti-spoofing+1Physiological-Physical Feature Fusion for Automatic Voice Spoofing Detection
Speaker verification systems have been used in many production scenarios in recent years. Unfortunately, they are still highly prone to different kinds of spoofing attacks such as voice conversion and speech synthesis, e…
Speaker VerificationSpeech SynthesisVoice ConversionSpoof detection using time-delay shallow neural network and feature switching
Detecting spoofed utterances is a fundamental problem in voice-based biometrics. Spoofing can be performed either by logical accesses like speech synthesis, voice conversion or by physical accesses such as replaying the …
Speaker VerificationSpeech SynthesisVoice Anti-spoofingVoice ConversionAdversarial Transformation of Spoofing Attacks for Voice Biometrics
Voice biometric systems based on automatic speaker verification (ASV) are exposed to \textit{spoofing} attacks which may compromise their security. To increase the robustness against such attacks, anti-spoofing or presen…
Speaker VerificationVoice ConversionVoiceCoach: Interactive Evidence-based Training for Voice Modulation Skills in Public Speaking
The modulation of voice properties, such as pitch, volume, and speed, is crucial for delivering a successful public speech. However, it is challenging to master different voice modulation skills. Though many guidelines a…
Sentence