paper-with-me

Papers

How does end-to-end speech recognition training impact speech enhancement artifacts?

2023-11-20 · Kazuma Iwamoto, Tsubasa Ochiai, Marc Delcroix, Rintaro Ikeshita, Hiroshi Sato, Shoko Araki, Shigeru Katagiri

Jointly training a speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end has been investigated as a way to mitigate the influence of \emph{processing distortion} generated by single-channel SE on ASR. In this paper, we investigate the effect of such joint training on the signal-level characteristics of the enhanced signals from the viewpoint of the decomposed noise and artifact errors. The experimental analyses provide two novel findings: 1) ASR-level training of the SE front-end reduces the artifact errors while increasing the noise errors, and 2) simply interpolating the enhanced and observed signals, which achieves a similar effect of reducing artifacts and increasing noise, improves ASR performance without jointly modifying the SE and ASR modules, even for a strong ASR back-end using a WavLM feature extractor. Our findings provide a better understanding of the effect of joint training and a novel insight for designing an ASR agnostic SE front-end.

📄 PDF Abstract BibTeX arXiv:2311.11599

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Investigating the Lombard Effect Influence on End-to-End Audio-Visual Speech Recognition

2019-06-05 · Pingchuan Ma, Stavros Petridis, Maja Pantic

Several audio-visual speech recognition models have been recently proposed which aim to improve the robustness over audio-only models in the presence of noise. However, almost all of them ignore the impact of the Lombard…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?

2022-04-27 · Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu 외

Recently, self-supervised learning (SSL) has demonstrated strong performance in speaker recognition, even if the pre-training objective is designed for speech recognition. In this paper, we study which factor leads to th…

Self-Supervised LearningSpeaker RecognitionSpeaker Verificationspeech-recognition+1

Measuring the Impact of Individual Domain Factors in Self-Supervised Pre-Training

2022-03-01 · Ramon Sanabria, Wei-Ning Hsu, Alexei Baevski, Michael Auli

Human speech data comprises a rich set of domain factors such as accent, syntactic and semantic variety, or acoustic environment. Previous work explores the effect of domain mismatch in automatic speech recognition betwe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A Study of Gender Impact in Self-supervised Models for Speech-to-Text Systems

2022-04-04 · Marcely Zanon Boito, Laurent Besacier, Natalia Tomashenko, Yannick Estève

Self-supervised models for speech processing emerged recently as popular foundation blocks in speech processing pipelines. These models are pre-trained on unlabeled audio data and then used in speech processing downstrea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Fairnessspeech-recognition+2

Back Transcription as a Method for Evaluating Robustness of Natural Language Understanding Models to Speech Recognition Errors

2023-10-25 · Marek Kubis, Paweł Skórzewski, Marcin Sowański, Tomasz Ziętkiewicz

In a spoken dialogue system, an NLU model is preceded by a speech recognition system that can deteriorate the performance of natural language understanding. This paper proposes a method for investigating the impact of sp…

en-US domain classificationen-US Intent Classificationen-US Slot FillingNatural Language Understanding+4