paper-with-me

홈 › Papers

Adapting TTS models For New Speakers using Transfer Learning

2021-10-12 · Paarth Neekhara, Jason Li, Boris Ginsburg

Training neural text-to-speech (TTS) models for a new speaker typically requires several hours of high quality speech data. Prior works on voice cloning attempt to address this challenge by adapting pre-trained multi-speaker TTS models for a new voice, using a few minutes of speech data of the new speaker. However, publicly available large multi-speaker datasets are often noisy, thereby resulting in TTS models that are not suitable for use in products. We address this challenge by proposing transfer-learning guidelines for adapting high quality single-speaker TTS models for a new speaker, using only a few minutes of speech data. We conduct an extensive study using different amounts of data for a new speaker and evaluate the synthesized speech in terms of naturalness and voice/style similarity to the target speaker. We find that fine-tuning a single-speaker TTS model on just 30 minutes of data, can yield comparable performance to a model trained from scratch on more than 27 hours of data for both male and female target speakers.

📄 PDF Abstract BibTeX arXiv:2110.05798

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to SpeechTransfer LearningVoice Cloning

Similar Papers 제목 키워드 기반

StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

2024-12-06 · Jixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning 외

Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using langu…

Voice Conversion

Adapting Multi-Lingual ASR Models for Handling Multiple Talkers

2023-05-30 · Chenda Li, Yao Qian, Zhuo Chen, Naoyuki Kanda 외

State-of-the-art large-scale universal speech models (USMs) show a decent automatic speech recognition (ASR) performance across multiple domains and languages. However, it remains a challenge for these models to recogniz…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Unsupervised Personalization of an Emotion Recognition System: The Unique Properties of the Externalization of Valence in Speech

2022-01-19 · Kusha Sridhar, Carlos Busso

The prediction of valence from speech is an important, but challenging problem. The externalization of valence in speech has speaker-dependent cues, which contribute to performances that are often significantly lower tha…

Emotion RecognitionPredictionSpeech Emotion RecognitionTransfer Learning

Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language

2024-09-02 · Jeong Hun Yeo, Chae Won Kim, Hyunjun Kim, Hyeongseop Rha 외

Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variat…

Lip ReadingSentence

Learning to adapt: a meta-learning approach for speaker adaptation

2018-08-30 · Ondřej Klejch, Joachim Fainberg, Peter Bell

The performance of automatic speech recognition systems can be improved by adapting an acoustic model to compensate for the mismatch between training and testing conditions, for example by adapting to unseen speakers. Th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Meta-Learningspeech-recognition+1