paper-with-me

홈 › Papers

Supervector Compression Strategies to Speed up I-Vector System Development

2018-05-03 · Ville Vestman, Tomi Kinnunen

The front-end factor analysis (FEFA), an extension of principal component analysis (PPCA) tailored to be used with Gaussian mixture models (GMMs), is currently the prevalent approach to extract compact utterance-level features (i-vectors) for automatic speaker verification (ASV) systems. Little research has been conducted comparing FEFA to the conventional PPCA applied to maximum a posteriori (MAP) adapted GMM supervectors. We study several alternative methods, including PPCA, factor analysis (FA), and two supervised approaches, supervised PPCA (SPPCA) and the recently proposed probabilistic partial least squares (PPLS), to compress MAP-adapted GMM supervectors. The resulting i-vectors are used in ASV tasks with a probabilistic linear discriminant analysis (PLDA) back-end. We experiment on two different datasets, on the telephone condition of NIST SRE 2010 and on the recent VoxCeleb corpus collected from YouTube videos containing celebrity interviews recorded in various acoustical and technical conditions. The results suggest that, in terms of ASV accuracy, the supervector compression approaches are on a par with FEFA. The supervised approaches did not result in improved performance. In comparison to FEFA, we obtained more than hundred-fold (100x) speedups in the total variability model (TVM) training using the PPCA and FA supervector compression approaches.

📄 PDF Abstract BibTeX arXiv:1805.01156

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

Differentiable Supervector Extraction for Encoding Speaker and Phrase Information in Text Dependent Speaker Verification

2018-12-22 · Victoria Mingote, Antonio Miguel, Alfonso Ortega, Eduardo Lleida

In this paper, we propose a new differentiable neural network alignment mechanism for text-dependent speaker verification which uses alignment models to produce a supervector representation of an utterance. Unlike previo…

Speaker VerificationText-Dependent Speaker Verification

Optimization of the Area Under the ROC Curve using Neural Network Supervectors for Text-Dependent Speaker Verification

2019-01-31 · Victoria Mingote, Antonio Miguel, Alfonso Ortega, Eduardo Lleida

This paper explores two techniques to improve the performance of text-dependent speaker verification systems based on deep neural networks. Firstly, we propose a general alignment mechanism to keep the temporal structure…

Speaker VerificationText-Dependent Speaker VerificationTriplet

Multilayer bootstrap network for unsupervised speaker recognition

2015-09-21 · Xiao-Lei Zhang

We apply multilayer bootstrap network (MBN), a recent proposed unsupervised learning method, to unsupervised speaker recognition. The proposed method first extracts supervectors from an unsupervised universal background …

ClusteringSpeaker Recognition

Identification/Segmentation of Indian Regional Languages with Singular Value Decomposition based Feature Embedding

2020-05-17

language identification (LID) is identifing a language in a given spoken utterance. Language segmentation is equally inportant as language identification where language boundaries can be spotted in a multi language utter…

Language IdentificationSegmentation

Multimodal Sparse Coding for Event Detection

2016-05-17 · Youngjune Gwon, William Campbell, Kevin Brady, Douglas Sturim 외

Unsupervised feature learning methods have proven effective for classification tasks based on a single modality. We present multimodal sparse coding for learning feature representations shared across multiple modalities.…

ClassificationEvent DetectionGeneral Classification