paper-with-me

Papers

Optimized Power Normalized Cepstral Coefficients towards Robust Deep Speaker Verification

2021-09-24 · Xuechen Liu, Md Sahidullah, Tomi Kinnunen

After their introduction to robust speech recognition, power normalized cepstral coefficient (PNCC) features were successfully adopted to other tasks, including speaker verification. However, as a feature extractor with long-term operations on the power spectrogram, its temporal processing and amplitude scaling steps dedicated on environmental compensation may be redundant. Further, they might suppress intrinsic speaker variations that are useful for speaker verification based on deep neural networks (DNN). Therefore, in this study, we revisit and optimize PNCCs by ablating its medium-time processor and by introducing channel energy normalization. Experimental results with a DNN-based speaker verification system indicate substantial improvement over baseline PNCCs on both in-domain and cross-domain scenarios, reflected by relatively 5.8% and 61.2% maximum lower equal error rate on VoxCeleb1 and VoxMovies, respectively.

📄 PDF Abstract BibTeX arXiv:2109.12058

Code (0)

등록된 구현이 없습니다.

Tasks

Robust Speech RecognitionSpeaker Verificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A Study of Acoustic Features in Arabic Speaker Identification under Noisy Environmental Conditions

2021-10-23 · Zhor Benhafid, Kawthar Yasmine Zergat, Abderrahmane Amrouche

One of the major parts of the voice recognition field is the choice of acoustic features which have to be robust against the variability of the speech signal, mismatched conditions, and noisy environments. Thus, differen…

Speaker Identification

A Comparison of Features for Replay Attack Detection

2019-12-12 · IOP Conf. Series: Journal of Physics: Conf. Series 1229 2019 12 · Zhifeng Xiea, Weibin Zhangb, Zhuxin Chen and Xiangmin Xu

Speaker verification (ASV) systems are still vulnerable to different kinds of spoofing attacks, especially replay attack due to high-quality playback devices. Many countermeasures have been developed recently. Most of th…

Speaker Verification

DNN Filter Bank Cepstral Coefficients for Spoofing Detection

2017-02-13 · Hong Yu, Zheng-Hua Tan, Zhanyu Ma, Jun Guo

With the development of speech synthesis techniques, automatic speaker verification systems face the serious challenge of spoofing attack. In order to improve the reliability of speaker verification systems, we develop a…

Speaker VerificationSpeech Synthesis

Pronunciation recognition of English phonemes /\textipa{@}/, /æ/, /\textipa{A}:/ and /\textipa{2}/ using Formants and Mel Frequency Cepstral Coefficients

2017-02-23 · Keith Y. Patarroyo, Vladimir Vargas-Calderón

The Vocal Joystick Vowel Corpus, by Washington University, was used to study monophthongs pronounced by native English speakers. The objective of this study was to quantitatively measure the extent at which speech recogn…

speech-recognitionSpeech Recognition

Optimization of data-driven filterbank for automatic speaker verification

2020-07-21 · Susanta Sarangi, Md Sahidullah, Goutam Saha

Most of the speech processing applications use triangular filters spaced in mel-scale for feature extraction. In this paper, we propose a new data-driven filter design method which optimizes filter parameters from a give…

Speaker Verification