Data Generation Using Pass-phrase-dependent Deep Auto-encoders for Text-Dependent Speaker Verification
In this paper, we propose a novel method that trains pass-phrase specific deep neural network (PP-DNN) based auto-encoders for creating augmented data for text-dependent speaker verification (TD-SV). Each PP-DNN auto-encoder is trained using the utterances of a particular pass-phrase available in the target enrollment set with two methods: (i) transfer learning and (ii) training from scratch. Next, feature vectors of a given utterance are fed to the PP-DNNs and the output from each PP-DNN at frame-level is considered one new set of generated data. The generated data from each PP-DNN is then used for building a TD-SV system in contrast to the conventional method that considers only the evaluation data available. The proposed approach can be considered as the transformation of data to the pass-phrase specific space using a non-linear transformation learned by each PP-DNN. The method develops several TD-SV systems with the number equal to the number of PP-DNNs separately trained for each pass-phrases for the evaluation. Finally, the scores of the different TD-SV systems are fused for decision making. Experiments are conducted on the RedDots challenge 2016 database for TD-SV using short utterances. Results show that the proposed method improves the performance for both conventional cepstral feature and deep bottleneck feature using both Gaussian mixture model - universal background model (GMM-UBM) and i-vector framework.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingSpeaker VerificationText-Dependent Speaker VerificationTransfer LearningSimilar Papers 제목 키워드 기반
Retrieval-Augmented Multilingual Keyphrase Generation with Retriever-Generator Iterative Training
Keyphrase generation is the task of automatically predicting keyphrases given a piece of long text. Despite its recent flourishing, keyphrase generation on non-English languages haven't been vastly investigated. In this …
Keyphrase GenerationPassage RetrievalRetrievalSpoken Pass-Phrase Verification in the i-vector Space
The task of spoken pass-phrase verification is to decide whether a test utterance contains the same phrase as given enrollment utterances. Beside other applications, pass-phrase verification can complement an independent…
Speaker VerificationText-Dependent Speaker VerificationIncorporating Pass-Phrase Dependent Background Models for Text-Dependent Speaker Verification
In this paper, we propose pass-phrase dependent background models (PBMs) for text-dependent (TD) speaker verification (SV) to integrate the pass-phrase identification process into the conventional TD-SV system, where a P…
Speaker VerificationText-Dependent Speaker VerificationTime-Contrastive Learning Based Deep Bottleneck Features for Text-Dependent Speaker Verification
There are a number of studies about extraction of bottleneck (BN) features from deep neural networks (DNNs)trained to discriminate speakers, pass-phrases and triphone states for improving the performance of text-dependen…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringContrastive Learning+4Time-Contrastive Learning Based DNN Bottleneck Features for Text-Dependent Speaker Verification
In this paper, we present a time-contrastive learning (TCL) based bottleneck (BN)feature extraction method for speech signals with an application to text-dependent (TD) speaker verification (SV). It is well-known that sp…
Contrastive LearningSpeaker VerificationText-Dependent Speaker Verification