paper-with-me

Papers

Deep multi-metric learning for text-independent speaker verification

2020-07-17 · Jiwei Xu, Xinggang Wang, Bin Feng, Wenyu Liu

Text-independent speaker verification is an important artificial intelligence problem that has a wide spectrum of applications, such as criminal investigation, payment certification, and interest-based customer services. The purpose of text-independent speaker verification is to determine whether two given uncontrolled utterances originate from the same speaker or not. Extracting speech features for each speaker using deep neural networks is a promising direction to explore and a straightforward solution is to train the discriminative feature extraction network by using a metric learning loss function. However, a single loss function often has certain limitations. Thus, we use deep multi-metric learning to address the problem and introduce three different losses for this problem, i.e., triplet loss, n-pair loss and angular loss. The three loss functions work in a cooperative way to train a feature extraction network equipped with Residual connections and squeeze-and-excitation attention. We conduct experiments on the large-scale \texttt{VoxCeleb2} dataset, which contains over a million utterances from over $6,000$ speakers, and the proposed deep neural network obtains an equal error rate of $3.48\%$, which is a very competitive result. Codes for both training and testing and pretrained models are available at \url{https://github.com/GreatJiweix/DmmlTiSV}, which is the first publicly available code repository for large-scale text-independent speaker verification with performance on par with the state-of-the-art systems.

📄 PDF Abstract BibTeX arXiv:2007.10479

Code (1)

GreatJiweix/DmmlTiSV 공식 구현 pytorch

Tasks

Metric LearningSpeaker VerificationText-Independent Speaker VerificationTriplet

Similar Papers 제목 키워드 기반

Few Shot Text-Independent speaker verification using 3D-CNN

2020-08-25 · Prateek Mishra

Facial recognition system is one of the major successes of Artificial intelligence and has been used a lot over the last years. But, images are not the only biometric present: audio is another possible biometric that can…

Speaker VerificationText-Independent Speaker Verification

Multi-task Metric Learning for Text-independent Speaker Verification

2020-10-21 · Yafeng Chen, Wu Guo, Jingjing Shi, Jiajun Qi 외

In this work, we introduce metric learning (ML) to enhance the deep embedding learning for text-independent speaker verification (SV). Specifically, the deep speaker embedding network is trained with conventional cross e…

Metric LearningSpeaker VerificationText-Independent Speaker Verification

A Multi Purpose and Large Scale Speech Corpus in Persian and English for Speaker and Speech Recognition: the DeepMine Database

2019-12-08 · Hossein Zeinali, Lukáš Burget, Jan "Honza'' Černocký

DeepMine is a speech database in Persian and English designed to build and evaluate text-dependent, text-prompted, and text-independent speaker verification, as well as Persian speech recognition systems. It contains mor…

Speaker Verificationspeech-recognitionSpeech RecognitionText-Dependent Speaker Verification+1

Deep Speaker Vectors for Semi Text-independent Speaker Verification

2015-05-24 · Lantian Li, Dong Wang, Zhiyong Zhang, Thomas Fang Zheng

Recent research shows that deep neural networks (DNNs) can be used to extract deep speaker vectors (d-vectors) that preserve speaker characteristics and can be used in speaker verification. This new method has been teste…

Speaker RecognitionSpeaker VerificationText-Dependent Speaker VerificationText-Independent Speaker Recognition+1

An End-to-End Text-independent Speaker Verification Framework with a Keyword Adversarial Network

2019-08-06 · Sungrack Yun, Janghoon Cho, Jungyun Eum, Wonil Chang 외

This paper presents an end-to-end text-independent speaker verification framework by jointly considering the speaker embedding (SE) network and automatic speech recognition (ASR) network. The SE network learns to output …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Verificationspeech-recognition+3