paper-with-me

홈 › Papers

The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance

2022-04-11 · Lin Zhang, Xin Wang, Erica Cooper, Nicholas Evans, Junichi Yamagishi

Automatic speaker verification is susceptible to various manipulations and spoofing, such as text-to-speech synthesis, voice conversion, replay, tampering, adversarial attacks, and so on. We consider a new spoofing scenario called "Partial Spoof" (PS) in which synthesized or transformed speech segments are embedded into a bona fide utterance. While existing countermeasures (CMs) can detect fully spoofed utterances, there is a need for their adaptation or extension to the PS scenario. We propose various improvements to construct a significantly more accurate CM that can detect and locate short-generated spoofed speech segments at finer temporal resolutions. First, we introduce newly developed self-supervised pre-trained models as enhanced feature extractors. Second, we extend our PartialSpoof database by adding segment labels for various temporal resolutions. Since the short spoofed speech segments to be embedded by attackers are of variable length, six different temporal resolutions are considered, ranging from as short as 20 ms to as large as 640 ms. Third, we propose a new CM that enables the simultaneous use of the segment-level labels at different temporal resolutions as well as utterance-level labels to execute utterance- and segment-level detection at the same time. We also show that the proposed CM is capable of detecting spoofing at the utterance level with low error rates in the PS scenario as well as in a related logical access (LA) scenario. The equal error rates of utterance-level detection on the PartialSpoof database and ASVspoof 2019 LA database were 0.77 and 0.90%, respectively.

📄 PDF Abstract BibTeX arXiv:2204.05177

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationSpeech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisVoice Conversion

Similar Papers 제목 키워드 기반

LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation

2024-09-23 · Hieu-Thi Luong, Haoyang Li, Lin Zhang, Kong Aik Lee 외

Previous fake speech datasets were constructed from a defender's perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we creat…

Language ModelingLanguage ModellingLarge Language Modeltext-to-speech+2

An Initial Investigation for Detecting Partially Spoofed Audio

2021-04-06 · Lin Zhang, Xin Wang, Erica Cooper, Junichi Yamagishi 외

All existing databases of spoofed speech contain attack data that is spoofed in its entirety. In practice, it is entirely plausible that successful attacks can be mounted with utterances that are only partially spoofed. …

Voice Anti-spoofing

ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection

2019-04-14

ASVspoof, now in its third edition, is a series of community-led challenges which promote the development of countermeasures to protect automatic speaker verification (ASV) from the threat of spoofing. Advances in the 20…

Speaker Verification

NE-PADD: Leveraging Named Entity Knowledge for Robust Partial Audio Deepfake Detection via Attention Aggregation

2025-09-04 · Huhong Xian, Rui Liu, Berrak Sisman, Haizhou Li arxiv

Different from traditional sentence-level audio deepfake detection (ADD), partial audio deepfake detection (PADD) requires frame-level positioning of the location of fake speech. While some progress has been made in this…

Audio Deepfake Detection

ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale

2024-08-16 · Xin Wang, Hector Delgado, Hemlata Tak, Jee-weon Jung 외

ASVspoof 5 is the fifth edition in a series of challenges that promote the study of speech spoofing and deepfake attacks, and the design of detection solutions. Compared to previous challenges, the ASVspoof 5 database is…

Face SwappingSpeaker Verification