paper-with-me

Papers

Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora

2024-09-17 · Francesco Nespoli, Daniel Barreda, Patrick A. Naylor

In recent years, automatic speech recognition (ASR) models greatly improved transcription performance both in clean, low noise, acoustic conditions and in reverberant environments. However, all these systems rely on the availability of hundreds of hours of labelled training data in specific acoustic conditions. When such a training dataset is not available, the performance of the system is heavily impacted. For example, this happens when a specific acoustic environment or a particular population of speakers is under-represented in the training dataset. Specifically, in this paper we investigate the effect of accented speech data on an off-the-shelf ASR system. Furthermore, we suggest a strategy based on zero-shot text-to-speech to augment the accented speech corpora. We show that this augmentation method is able to mitigate the loss in performance of the ASR system on accented data up to 5% word error rate reduction (WERR). In conclusion, we demonstrate that by incorporating a modest fraction of real with synthetically generated data, the ASR system exhibits superior performance compared to a model trained exclusively on authentic accented speech with up to 14% WERR.

📄 PDF Abstract BibTeX arXiv:2409.11107

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognitiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM

2024-11-20 · Jiawei Yu, Yuang Li, Xiaosong Qiao, Huan Zhao 외

Text-to-speech (TTS) models have been widely adopted to enhance automatic speech recognition (ASR) systems using text-only corpora, thereby reducing the cost of labeling real speech data. Existing research primarily util…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+3

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

2026-06-18 · Satwinder Singh, Qianli Wang, Zihan Zhong, Clarion Mendes 외 arxiv

Automatic speech recognition remains unreliable for dysarthric speech due to data scarcity and high inter-speaker variability. While synthetic data can address these gaps, traditional methods often require extensive spea…

Speech RecognitionData Augmentation

Speech collage: code-switched audio generation by collaging monolingual corpora

2023-09-27 · Amir Hussein, Dorsa Zeinali, Ondřej Klejch, Matthew Wiesner 외

Designing effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a …

Audio GenerationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1

Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement

2025-01-23 · Jae-Sung Bae, Anastasia Kuznetsova, Dinesh Manocha, John Hershey 외

This paper presents a new challenge that calls for zero-shot text-to-speech (TTS) systems to augment speech data for the downstream task, personalized speech enhancement (PSE), as part of the Generative Data Augmentation…

Data AugmentationSpeech EnhancementSpeech SynthesisSynthetic Data Generation+2

ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis

2026-03-04 · Youngwon Choi, Jinwoo Oh, Hwayeon Kim, Hyeonyu Kim arxiv

We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically dive…

Data AugmentationSpeech Synthesis