paper-with-me

Papers

Speaker Recognition Based on Deep Learning: An Overview

2020-12-02 · Zhongxin Bai, Xiao-Lei Zhang

Speaker recognition is a task of identifying persons from their voices. Recently, deep learning has dramatically revolutionized speaker recognition. However, there is lack of comprehensive reviews on the exciting progress. In this paper, we review several major subtasks of speaker recognition, including speaker verification, identification, diarization, and robust speaker recognition, with a focus on deep-learning-based methods. Because the major advantage of deep learning over conventional methods is its representation ability, which is able to produce highly abstract embedding features from utterances, we first pay close attention to deep-learning-based speaker feature extraction, including the inputs, network structures, temporal pooling strategies, and objective functions respectively, which are the fundamental components of many speaker recognition subtasks. Then, we make an overview of speaker diarization, with an emphasis of recent supervised, end-to-end, and online diarization. Finally, we survey robust speaker recognition from the perspectives of domain adaptation and speech enhancement, which are two major approaches of dealing with domain mismatch and noise problems. Popular and recently released corpora are listed at the end of the paper.

📄 PDF Abstract BibTeX arXiv:2012.00931

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningDomain Adaptationspeaker-diarizationSpeaker DiarizationSpeaker RecognitionSpeaker VerificationSpeech Enhancement

Similar Papers 제목 키워드 기반

Adaptation Algorithms for Neural Network-Based Speech Recognition: An Overview

2020-08-14 · Peter Bell, Joachim Fainberg, Ondrej Klejch, Jinyu Li 외

We present a structured overview of adaptation algorithms for neural network-based speech recognition, considering both hybrid hidden Markov model / neural network systems and end-to-end neural network systems, with a fo…

Data AugmentationDomain Adaptationspeech-recognitionSpeech Recognition

Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning

2024-07-21 · Shuai Wang, Zhengyang Chen, Kong Aik Lee, Yanmin Qian 외

Speaker individuality information is among the most critical elements within speech signals. By thoroughly and accurately modeling this information, it can be utilized in various intelligent speech applications, such as …

Representation LearningSelf-Supervised Learningspeaker-diarizationSpeaker Diarization+3

State-of-the-art in speaker recognition

2022-02-23 · Marcos Faundez-Zanuy, Enric Monte-Moreno

Recent advances in speech technologies have produced new tools that can be used to improve the performance and flexibility of speaker recognition While there are few degrees of freedom or alternative methods when using f…

Speaker Recognition

A Survey on Paralinguistics in Tamil Speech Processing

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Anosha Ignatius, Uthayasanker Thayasivam

Speech carries not only the semantic content but also the paralinguistic information which captures the speaking style. Speaker traits and emotional states affect how words are being spoken. The research on paralinguisti…

Emotion RecognitionSpeaker Identificationspeech-recognitionSpeech Recognition+1

End-to-end training of time domain audio separation and recognition

2019-12-18 · Thilo von Neumann, Keisuke Kinoshita, Lukas Drude, Christoph Boeddeker 외

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based t…

Speaker Recognitionspeech-recognitionSpeech RecognitionSpeech Separation