paper-with-me

홈 › Papers

Differentiable Supervector Extraction for Encoding Speaker and Phrase Information in Text Dependent Speaker Verification

2018-12-22 · Victoria Mingote, Antonio Miguel, Alfonso Ortega, Eduardo Lleida

In this paper, we propose a new differentiable neural network alignment mechanism for text-dependent speaker verification which uses alignment models to produce a supervector representation of an utterance. Unlike previous works with similar approaches, we do not extract the embedding of an utterance from the mean reduction of the temporal dimension. Our system replaces the mean by a phrase alignment model to keep the temporal structure of each phrase which is relevant in this application since the phonetic information is part of the identity in the verification task. Moreover, we can apply a convolutional neural network as front-end, and thanks to the alignment process being differentiable, we can train the whole network to produce a supervector for each utterance which will be discriminative with respect to the speaker and the phrase simultaneously. As we show, this choice has the advantage that the supervector encodes the phrase and speaker information providing good performance in text-dependent speaker verification tasks. In this work, the process of verification is performed using a basic similarity metric, due to simplicity, compared to other more elaborate models that are commonly used. The new model using alignment to produce supervectors was tested on the RSR2015-Part I database for text-dependent speaker verification, providing competitive results compared to similar size networks using the mean to extract embeddings.

📄 PDF Abstract BibTeX arXiv:1812.09484

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationText-Dependent Speaker Verification

Similar Papers 제목 키워드 기반

Optimization of the Area Under the ROC Curve using Neural Network Supervectors for Text-Dependent Speaker Verification

2019-01-31 · Victoria Mingote, Antonio Miguel, Alfonso Ortega, Eduardo Lleida

This paper explores two techniques to improve the performance of text-dependent speaker verification systems based on deep neural networks. Firstly, we propose a general alignment mechanism to keep the temporal structure…

Speaker VerificationText-Dependent Speaker VerificationTriplet

Multilayer bootstrap network for unsupervised speaker recognition

2015-09-21 · Xiao-Lei Zhang

We apply multilayer bootstrap network (MBN), a recent proposed unsupervised learning method, to unsupervised speaker recognition. The proposed method first extracts supervectors from an unsupervised universal background …

ClusteringSpeaker Recognition

Supervector Compression Strategies to Speed up I-Vector System Development

2018-05-03 · Ville Vestman, Tomi Kinnunen

The front-end factor analysis (FEFA), an extension of principal component analysis (PPCA) tailored to be used with Gaussian mixture models (GMMs), is currently the prevalent approach to extract compact utterance-level fe…

Speaker Verification

Improved Frame Level Features and SVM Supervectors Approach for the Recogniton of Emotional States from Speech: Application to categorical and dimensional states

2014-06-23 · Imen Trabelsi, Dorra Ben Ayed, Noureddine Ellouze

The purpose of speech emotion recognition system is to classify speakers utterances into different emotional states such as disgust, boredom, sadness, neutral and happiness. Speech features that are commonly used in spee…

Emotion RecognitionSpeech Emotion Recognition

Time-Contrastive Learning Based DNN Bottleneck Features for Text-Dependent Speaker Verification

2017-04-06 · Achintya Kr. Sarkar, Zheng-Hua Tan

In this paper, we present a time-contrastive learning (TCL) based bottleneck (BN)feature extraction method for speech signals with an application to text-dependent (TD) speaker verification (SV). It is well-known that sp…

Contrastive LearningSpeaker VerificationText-Dependent Speaker Verification