paper-with-me

홈 › Papers

Analysis of DNN Speech Signal Enhancement for Robust Speaker Recognition

2018-11-19

In this work, we present an analysis of a DNN-based autoencoder for speech enhancement, dereverberation and denoising. The target application is a robust speaker verification (SV) system. We start our approach by carefully designing a data augmentation process to cover wide range of acoustic conditions and obtain rich training data for various components of our SV system. We augment several well-known databases used in SV with artificially noised and reverberated data and we use them to train a denoising autoencoder (mapping noisy and reverberated speech to its clean version) as well as an x-vector extractor which is currently considered as state-of-the-art in SV. Later, we use the autoencoder as a preprocessing step for text-independent SV system. We compare results achieved with autoencoder enhancement, multi-condition PLDA training and their simultaneous use. We present a detailed analysis with various conditions of NIST SRE 2010, 2016, PRISM and with re-transmitted data. We conclude that the proposed preprocessing can significantly improve both i-vector and x-vector baselines and that this technique can be used to build a robust SV system for various target domains.

📄 PDF Abstract BibTeX arXiv:1811.07629

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDenoisingSpeaker RecognitionSpeaker VerificationSpeech Enhancement

Similar Papers 제목 키워드 기반

Speaker Re-identification with Speaker Dependent Speech Enhancement

2020-05-15 · Yanpei Shi, Qiang Huang, Thomas Hain

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditiona…

Speaker RecognitionSpeech Enhancement

Robust Speaker Recognition Using Speech Enhancement And Attention Model

2020-01-14 · Yanpei Shi, Qiang Huang, Thomas Hain

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performance when speech signals are corrupted by n…

Speaker IdentificationSpeaker RecognitionSpeech Enhancement

Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention

2020-02-14 · Yuma Koizumi, Kohei Yatabe, Marc Delcroix, Yoshiki Masuyama 외

This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studi…

Multi-Task LearningSpeaker IdentificationSpeech Enhancementspeech-recognition+1

A Conformer-based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech Separation

2021-11-18 · Tom O'Malley, Arun Narayanan, Quan Wang, Alex Park 외

We present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speech enhancement, and speech separation. Th…

Acoustic echo cancellationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancement+3

Streaming Noise Context Aware Enhancement For Automatic Speech Recognition in Multi-Talker Environments

2022-05-17 · Joe Caroselli, Arun Narayanan, Yiteng Huang

One of the most challenging scenarios for smart speakers is multi-talker, when target speech from the desired speaker is mixed with interfering speech from one or more speakers. A smart assistant needs to determine which…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1