Deep learning methods in speaker recognition: a review
This paper summarizes the applied deep learning practices in the field of speaker recognition, both verification and identification. Speaker recognition has been a widely used field topic of speech technology. Many research works have been carried out and little progress has been achieved in the past 5-6 years. However, as deep learning techniques do advance in most machine learning fields, the former state-of-the-art methods are getting replaced by them in speaker recognition too. It seems that DL becomes the now state-of-the-art solution for both speaker verification and identification. The standard x-vectors, additional to i-vectors, are used as baseline in most of the novel works. The increasing amount of gathered data opens up the territory to DL, where they are the most effective.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningSpeaker RecognitionSpeaker VerificationSimilar Papers 제목 키워드 기반
Speaker Recognition Based on Deep Learning: An Overview
Speaker recognition is a task of identifying persons from their voices. Recently, deep learning has dramatically revolutionized speaker recognition. However, there is lack of comprehensive reviews on the exciting progres…
Deep LearningDomain Adaptationspeaker-diarizationSpeaker Diarization+3A Review of Speaker Diarization: Recent Advances with Deep Learning
Speaker diarization is a task to label audio or video recordings with classes that correspond to speaker identity, or in short, a task to identify "who spoke when". In the early years, speaker diarization algorithms were…
Deep LearningRetrievalspeaker-diarizationSpeaker Diarization+2Utterance partitioning for speaker recognition: an experimental review and analysis with new findings under GMM-SVM framework
The performance of speaker recognition system is highly dependent on the amount of speech used in enrollment and test. This work presents a detailed experimental review and analysis of the GMM-SVM based speaker recogniti…
Speaker RecognitionCNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
The first Chinese Continuous Visual Speech Recognition Challenge aimed to probe the performance of Large Vocabulary Continuous Visual Speech Recognition (LVC-VSR) on two tasks: (1) Single-speaker VSR for a particular spe…
speech-recognitionSpeech RecognitionVisual Speech RecognitionA Comprehensive Survey on Multi-modal Conversational Emotion Recognition with Deep Learning
Multi-modal conversation emotion recognition (MCER) aims to recognize and track the speaker's emotional state using text, speech, and visual information in the conversation scene. Analyzing and studying MCER issues is si…
Emotion Recognition