paper-with-me

Papers

Deep neural network based i-vector mapping for speaker verification using short utterances

2018-10-16 · Jinxi Guo, Ning Xu, Kailun Qian, Yang Shi, Kaiyuan Xu, Ying-Nian Wu, Abeer Alwan

Text-independent speaker recognition using short utterances is a highly challenging task due to the large variation and content mismatch between short utterances. I-vector based systems have become the standard in speaker verification applications, but they are less effective with short utterances. In this paper, we first compare two state-of-the-art universal background model training methods for i-vector modeling using full-length and short utterance evaluation tasks. The two methods are Gaussian mixture model (GMM) based and deep neural network (DNN) based methods. The results indicate that the I-vector_DNN system outperforms the I-vector_GMM system under various durations. However, the performances of both systems degrade significantly as the duration of the utterances decreases. To address this issue, we propose two novel nonlinear mapping methods which train DNN models to map the i-vectors extracted from short utterances to their corresponding long-utterance i-vectors. The mapped i-vector can restore missing information and reduce the variance of the original short-utterance i-vectors. The proposed methods both model the joint representation of short and long utterance i-vectors by using autoencoder. Experimental results using the NIST SRE 2010 dataset show that both methods provide significant improvement and result in a max of 28.43% relative improvement in Equal Error Rates from a baseline system, when using deep encoder with residual blocks and adding an additional phoneme vector. When further testing the best-validated models of SRE10 on the Speaker In The Wild dataset, the methods result in a 23.12% improvement on arbitrary-duration (1-5 s) short-utterance conditions.

📄 PDF Abstract BibTeX arXiv:1810.07309

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker RecognitionSpeaker VerificationText-Independent Speaker Recognition

Similar Papers 제목 키워드 기반

I-vector Transformation Using Conditional Generative Adversarial Networks for Short Utterance Speaker Verification

2018-04-01 · Jiacen Zhang, Nakamasa Inoue, Koichi Shinoda

I-vector based text-independent speaker verification (SV) systems often have poor performance with short utterances, as the biased phonetic distribution in a short utterance makes the extracted i-vector unreliable. This …

Generative Adversarial NetworkSpeaker VerificationText-Independent Speaker Verification

Quality Measures for Speaker Verification with Short Utterances

2019-01-29 · Arnab Poddar, Md Sahidullah, Goutam Saha

The performances of the automatic speaker verification (ASV) systems degrade due to the reduction in the amount of speech used for enrollment and verification. Combining multiple systems based on different features and c…

Speaker RecognitionSpeaker Verification

Deep Speaker Embeddings for Far-Field Speaker Recognition on Short Utterances

2020-02-14 · Aleksei Gusev, Vladimir Volokhov, Tseren Andzhukaev, Sergey Novoselov 외

Speaker recognition systems based on deep speaker embeddings have achieved significant performance in controlled conditions according to the results obtained for early NIST SRE (Speaker Recognition Evaluation) datasets. …

Speaker RecognitionSpeaker Verification

Segment Aggregation for short utterances speaker verification using raw waveforms

2020-05-07 · Seung-bin Kim, Jee-weon Jung, Hye-jin Shim, Ju-ho Kim 외

Most studies on speaker verification systems focus on long-duration utterances, which are composed of sufficient phonetic information. However, the performances of these systems are known to degrade when short-duration u…

Speaker Verification

End-to-End Attention based Text-Dependent Speaker Verification

2017-01-03 · Shi-Xiong Zhang, Zhuo Chen, Yong Zhao, Jinyu Li 외

A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetically discriminative/speaker discriminative DNNs as feature extractors for speaker verifica…

Speaker VerificationText-Dependent Speaker Verification