paper-with-me

Papers

Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?

2020-05-04 · Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Junichi Yamagishi

Previous work on speaker adaptation for end-to-end speech synthesis still falls short in speaker similarity. We investigate an orthogonal approach to the current speaker adaptation paradigms, speaker augmentation, by creating artificial speakers and by taking advantage of low-quality data. The base Tacotron2 model is modified to account for the channel and dialect factors inherent in these corpora. In addition, we describe a warm-start training strategy that we adopted for Tacotron2 training. A large-scale listening test is conducted, and a distance metric is adopted to evaluate synthesis of dialects. This is followed by an analysis on synthesis quality, speaker and dialect similarity, and a remark on the effectiveness of our speaker augmentation approach. Audio samples are available online.

📄 PDF Abstract BibTeX arXiv:2005.01245

Code (1)

nii-yamagishilab/multi-speaker-tacotron 공식 구현 tf

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Speaker verification-derived loss and data augmentation for DNN-based multispeaker speech synthesis

2021-06-03 · Beata Lorincz, Adriana Stan, Mircea Giurgiu

Building multispeaker neural network-based text-to-speech synthesis systems commonly relies on the availability of large amounts of high quality recordings from each speaker and conditioning the training process on the s…

Data AugmentationSpeaker VerificationSpeech Synthesistext-to-speech+2

ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion

2022-03-29 · Edresson Casanova, Christopher Shulby, Alexander Korolev, Arnaldo Candido Junior 외

We explore cross-lingual multi-speaker speech synthesis and cross-lingual voice conversion applied to data augmentation for automatic speech recognition (ASR) systems in low/medium-resource scenarios. Through extensive e…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+3

Cross-speaker style transfer for text-to-speech using data augmentation

2022-02-10 · Manuel Sam Ribeiro, Julian Roth, Giulia Comini, Goeric Huybrechts 외

We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting…

Data AugmentationStyle Transfertext-to-speechText to Speech+1

Relational Data Selection for Data Augmentation of Speaker-dependent Multi-band MelGAN Vocoder

2021-06-10 · Yi-Chiao Wu, Cheng-Hung Hu, Hung-Shin Lee, Yu-Huai Peng 외

Nowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available. Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impract…

Data AugmentationSpeaker Verification

Using Data Augmentations and VTLN to Reduce Bias in Dutch End-to-End Speech Recognition Systems

2023-07-05 · Tanvina Patel, Odette Scharenborg

Speech technology has improved greatly for norm speakers, i.e., adult native speakers of a language without speech impediments or strong accents. However, non-norm or diverse speaker groups show a distinct performance ga…

AnatomyData Augmentationspeech-recognitionSpeech Recognition