DNN Filter Bank Cepstral Coefficients for Spoofing Detection
With the development of speech synthesis techniques, automatic speaker verification systems face the serious challenge of spoofing attack. In order to improve the reliability of speaker verification systems, we develop a new filter bank based cepstral feature, deep neural network filter bank cepstral coefficients (DNN-FBCC), to distinguish between natural and spoofed speech. The deep neural network filter bank is automatically generated by training a filter bank neural network (FBNN) using natural and synthetic speech. By adding restrictions on the training rules, the learned weight matrix of FBNN is band-limited and sorted by frequency, similar to the normal filter bank. Unlike the manually designed filter bank, the learned filter bank has different filter shapes in different channels, which can capture the differences between natural and synthetic speech more effectively. The experimental results on the ASVspoof {2015} database show that the Gaussian mixture model maximum-likelihood (GMM-ML) classifier trained by the new feature performs better than the state-of-the-art linear frequency cepstral coefficients (LFCC) based classifier, especially on detecting unknown attacks.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker VerificationSpeech SynthesisSimilar Papers 제목 키워드 기반
A Comparison of Features for Replay Attack Detection
Speaker verification (ASV) systems are still vulnerable to different kinds of spoofing attacks, especially replay attack due to high-quality playback devices. Many countermeasures have been developed recently. Most of th…
Speaker VerificationCepstral Coefficients for Earthquake Damage Assessment of Bridges Leveraging Deep Learning
Bridges are indispensable elements in resilient communities as essential parts of the lifeline transportation systems. Knowledge about the functionality of bridge structures is crucial, especially after a major earthquak…
Optimization of data-driven filterbank for automatic speaker verification
Most of the speech processing applications use triangular filters spaced in mel-scale for feature extraction. In this paper, we propose a new data-driven filter design method which optimizes filter parameters from a give…
Speaker VerificationModified Mel Filter Bank to Compute MFCC of Subsampled Speech
Mel Frequency Cepstral Coefficients (MFCCs) are the most popularly used speech features in most speech and speaker recognition applications. In this work, we propose a modified Mel filter bank to extract MFCCs from subsa…
Speaker RecognitionAn explainability study of the constant Q cepstral coefficient spoofing countermeasure for automatic speaker verification
Anti-spoofing for automatic speaker verification is now a well established area of research, with three competitive challenges having been held in the last 6 years. A great deal of research effort over this time has been…
Speaker VerificationSpeech Synthesis