paper-with-me

Papers

Thech. Report: Genuinization of Speech waveform PMF for speaker detection spoofing and countermeasures

2023-10-09 · Itshak Lapidot, Jean-Francois Bonastre

In the context of spoofing attacks in speaker recognition systems, we observed that the waveform probability mass function (PMF) of genuine speech differs significantly from the PMF of speech resulting from the attacks. This is true for synthesized or converted speech as well as replayed speech. We also noticed that this observation seems to have a significant impact on spoofing detection performance. In this article, we propose an algorithm, denoted genuinization, capable of reducing the waveform distribution gap between authentic speech and spoofing speech. Our genuinization algorithm is evaluated on ASVspoof 2019 challenge datasets, using the baseline system provided by the challenge organization. We first assess the influence of genuinization on spoofing performance. Using genuinization for the spoofing attacks degrades spoofing detection performance by up to a factor of 10. Next, we integrate the genuinization algorithm in the spoofing countermeasures and we observe a huge spoofing detection improvement in different cases. The results of our experiments show clearly that waveform distribution plays an important role and must be taken into account by anti-spoofing systems.

📄 PDF Abstract BibTeX arXiv:2310.05534

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Recognition

Similar Papers 제목 키워드 기반

Speaker-independent raw waveform model for glottal excitation

2018-04-25 · Lauri Juvela, Vassilis Tsiaras, Bajibabu Bollepalli, Manu Airaksinen 외

Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated spe…

modelSpeech Synthesistext-to-speechText to Speech+2

Facetron: A Multi-speaker Face-to-Speech Model based on Cross-modal Latent Representations

2021-07-26 · Se-Yun Um, Jihyun Kim, Jihyun Lee, Hong-Goo Kang

In this paper, we propose a multi-speaker face-to-speech waveform generation model that also works for unseen speaker conditions. Using a generative adversarial network (GAN) with linguistic and speaker characteristic fe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Generative Adversarial NetworkLip Reading+2

Y-Vector: Multiscale Waveform Encoder for Speaker Embedding

2020-10-24 · Ge Zhu, Fei Jiang, Zhiyao Duan

State-of-the-art text-independent speaker verification systems typically use cepstral features or filter bank energies as speech features. Recent studies attempted to extract speaker embeddings directly from raw waveform…

Speaker VerificationText-Independent Speaker Verification

Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis

2020-11-10 · Erica Cooper, Xin Wang, Yi Zhao, Yusuke Yasuda 외

We explore pretraining strategies including choice of base corpus with the aim of choosing the best strategy for zero-shot multi-speaker end-to-end synthesis. We also examine choice of neural vocoder for waveform synthes…

Speech Synthesis

Universal MelGAN: A Robust Neural Vocoder for High-Fidelity Waveform Generation in Multiple Domains

2020-11-19 · Won Jang, Dan Lim, Jaesam Yoon

We propose Universal MelGAN, a vocoder that synthesizes high-fidelity speech in multiple domains. To preserve sound quality when the MelGAN-based structure is trained with a dataset of hundreds of speakers, we added mult…

text-to-speechText to Speech