paper-with-me

Papers

Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech Recognition

2022-05-09 · Catalin Zorila, Rama Doddipatla

Improving the accuracy of single-channel automatic speech recognition (ASR) in noisy conditions is challenging. Strong speech enhancement front-ends are available, however, they typically require that the ASR model is retrained to cope with the processing artifacts. In this paper we explore a speaker reinforcement strategy for improving recognition performance without retraining the acoustic model (AM). This is achieved by remixing the enhanced signal with the unprocessed input to alleviate the processing artifacts. We evaluate the proposed approach using a DNN speaker extraction based speech denoiser trained with a perceptually motivated loss function. Results show that (without AM retraining) our method yields about 23% and 25% relative accuracy gains compared with the unprocessed for the monoaural simulated and real CHiME-4 evaluation sets, respectively, and outperforms a state-of-the-art reference method.

📄 PDF Abstract BibTeX arXiv:2205.04433

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

AM 설명 없음

Similar Papers 제목 키워드 기반

Conditional Diffusion Model for Target Speaker Extraction

2023-10-07 · Theodor Nguyen, Guangzhi Sun, Xianrui Zheng, Chao Zhang 외

We propose DiffSpEx, a generative target speaker extraction method based on score-based generative modelling through stochastic differential equations. DiffSpEx deploys a continuous-time stochastic diffusion process in t…

modelTarget Speaker Extraction

Multi-stage Speaker Extraction with Utterance and Frame-Level Reference Signals

2020-11-19 · Meng Ge, Chenglin Xu, Longbiao Wang, Eng Siong Chng 외

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple…

USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction

2024-09-04 · Bang Zeng, Ming Li

Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, in which a speaker recogniti…

Speaker RecognitionSpeech SeparationTarget Speaker Extraction

Improving speaker discrimination of target speech extraction with time-domain SpeakerBeam

2020-01-23 · Marc Delcroix, Tsubasa Ochiai, Katerina Zmolikova, Keisuke Kinoshita 외

Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation u…

Speaker IdentificationSpeech ExtractionSpeech Separation

Informed Source Extraction With Application to Acoustic Echo Reduction

2020-11-09 · Mohamed Elminshawi, Wolfgang Mack, Emanuël A. P. Habets

Informed speaker extraction aims to extract a target speech signal from a mixture of sources given prior knowledge about the desired speaker. Recent deep learning-based methods leverage a speaker discriminative model tha…

Acoustic echo cancellationSpeaker Separation