paper-with-me

홈 › Papers

Low-Resource Text-to-Speech Using Specific Data and Noise Augmentation

2023-06-16 · Kishor Kayyar Lakshminarayana, Christian Dittmar, Nicola Pia, Emanuël Habets

Many neural text-to-speech architectures can synthesize nearly natural speech from text inputs. These architectures must be trained with tens of hours of annotated and high-quality speech data. Compiling such large databases for every new voice requires a lot of time and effort. In this paper, we describe a method to extend the popular Tacotron-2 architecture and its training with data augmentation to enable single-speaker synthesis using a limited amount of specific training data. In contrast to elaborate augmentation methods proposed in the literature, we use simple stationary noises for data augmentation. Our extension is easy to implement and adds almost no computational overhead during training and inference. Using only two hours of training data, our approach was rated by human listeners to be on par with the baseline Tacotron-2 trained with 23.5 hours of LJSpeech data. In addition, we tested our model with a semantically unpredictable sentences test, which showed that both models exhibit similar intelligibility levels.

📄 PDF Abstract BibTeX arXiv:2306.10152

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentationtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron

2025-01-10 · Kishor Kayyar Lakshminarayana, Frank Zalkow, Christian Dittmar, Nicola Pia 외

In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many di…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages

2026-04-09 · Lilian Wanzare, Cynthia Amol, Ezekiel Maina, Nelson Odhiambo 외 arxiv

AfriVoices-KE is a large-scale multilingual speech dataset comprising approximately 3,000 hours of audio across five Kenyan languages: Dholuo, Kikuyu, Kalenjin, Maasai, and Somali. The dataset includes 750 hours of scrip…

Speech Recognition

Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora

2024-09-17 · Francesco Nespoli, Daniel Barreda, Patrick A. Naylor

In recent years, automatic speech recognition (ASR) models greatly improved transcription performance both in clean, low noise, acoustic conditions and in reverberant environments. However, all these systems rely on the …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Monaural speech enhancement on drone via Adapter based transfer learning

2024-05-16 · Xingyu Chen, Hanwen Bi, Wei-Ting Lai, Fei Ma

Monaural Speech enhancement on drones is challenging because the ego-noise from the rotating motors and propellers leads to extremely low signal-to-noise ratios at onboard microphones. Although recent masking-based deep …

Speech EnhancementTransfer Learning

Noise Robust TTS for Low Resource Speakers using Pre-trained Model and Speech Enhancement

2020-05-26 · Dongyang Dai, Li Chen, Yu-Ping Wang, Mu Wang 외

With the popularity of deep neural network, speech synthesis task has achieved significant improvements based on the end-to-end encoder-decoder framework in the recent days. More and more applications relying on speech s…

DecoderSpeech EnhancementSpeech Synthesis