Optimizing Multi-Taper Features for Deep Speaker Verification
Multi-taper estimators provide low-variance power spectrum estimates that can be used in place of the windowed discrete Fourier transform (DFT) to extract speech features such as mel-frequency cepstral coefficients (MFCCs). Even if past work has reported promising automatic speaker verification (ASV) results with Gaussian mixture model-based classifiers, the performance of multi-taper MFCCs with deep ASV systems remains an open question. Instead of a static-taper design, we propose to optimize the multi-taper estimator jointly with a deep neural network trained for ASV tasks. With a maximum improvement on the SITW corpus of 25.8% in terms of equal error rate over the static-taper, our method helps preserve a balanced level of leakage and variance, providing more robustness.
Code (0)
등록된 구현이 없습니다.
Tasks
Open-Ended Question AnsweringSpeaker VerificationSimilar Papers 제목 키워드 기반
Application of Particle Swarm Optimization to Microwave Tapered Microstrip Lines
Application of metaheuristic algorithms has been of continued interest in the field of electrical engineering because of their powerful features. In this work special design is done for a tapered transmission line used f…
Electrical EngineeringMultiobjective OptimizationProsodic-Enhanced Siamese Convolutional Neural Networks for Cross-Device Text-Independent Speaker Verification
In this paper a novel cross-device text-independent speaker verification architecture is proposed. Majority of the state-of-the-art deep architectures that are used for speaker verification tasks consider Mel-frequency c…
Speaker VerificationText-Independent Speaker VerificationRobust End-to-End Speaker Verification Using EEG
In this paper we demonstrate that performance of a speaker verification system can be improved by concatenating electroencephalography (EEG) signal features with speech signal features or only using EEG signal features. …
EEGElectroencephalogram (EEG)Speaker VerificationTriplet Based Embedding Distance and Similarity Learning for Text-independent Speaker Verification
Speaker embeddings become growing popular in the text-independent speaker verification task. In this paper, we propose two improvements during the training stage. The improvements are both based on triplet cause the trai…
Speaker RecognitionSpeaker VerificationText-Independent Speaker VerificationTripletJoint Speaker Encoder and Neural Back-end Model for Fully End-to-End Automatic Speaker Verification with Multiple Enrollment Utterances
Conventional automatic speaker verification systems can usually be decomposed into a front-end model such as time delay neural network (TDNN) for extracting speaker embeddings and a back-end model such as statistics-base…
Data AugmentationSpeaker Verification