paper-with-me

Papers

Generalizable Zero-Shot Speaker Adaptive Speech Synthesis with Disentangled Representations

2023-08-24 · Wenbin Wang, Yang song, Sanjay Jha

While most research into speech synthesis has focused on synthesizing high-quality speech for in-dataset speakers, an equally essential yet unsolved problem is synthesizing speech for unseen speakers who are out-of-dataset with limited reference data, i.e., speaker adaptive speech synthesis. Many studies have proposed zero-shot speaker adaptive text-to-speech and voice conversion approaches aimed at this task. However, most current approaches suffer from the degradation of naturalness and speaker similarity when synthesizing speech for unseen speakers (i.e., speakers not in the training dataset) due to the poor generalizability of the model in out-of-distribution data. To address this problem, we propose GZS-TV, a generalizable zero-shot speaker adaptive text-to-speech and voice conversion model. GZS-TV introduces disentangled representation learning for both speaker embedding extraction and timbre transformation to improve model generalization and leverages the representation learning capability of the variational autoencoder to enhance the speaker encoder. Our experiments demonstrate that GZS-TV reduces performance degradation on unseen speakers and outperforms all baseline models in multiple datasets.

📄 PDF Abstract BibTeX arXiv:2308.13007

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSpeech Synthesistext-to-speechText to SpeechVoice Conversion

Similar Papers 제목 키워드 기반

AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios

2022-04-01 · Yihan Wu, Xu Tan, Bohan Li, Lei He 외

Adaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen spe…

Speech Synthesistext-to-speechText to Speech

GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech

2022-05-15 · Rongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui 외

Style transfer for out-of-domain (OOD) speech synthesis aims to generate speech samples with unseen style (e.g., speaker identity, emotion, and prosody) derived from an acoustic reference, while facing the following chal…

Speech SynthesisStyle Transfertext-to-speechText to Speech+1

USAT: A Universal Speaker-Adaptive Text-to-Speech Approach

2024-04-28 · Wenbin Wang, Yang song, Sanjay Jha

Conventional text-to-speech (TTS) research has predominantly focused on enhancing the quality of synthesized speech for speakers in the training dataset. The challenge of synthesizing lifelike speech for unseen, out-of-d…

Decodertext-to-speechText to Speech

ZET-Speech: Zero-shot adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models

2023-05-23 · Minki Kang, Wooseok Han, Sung Ju Hwang, Eunho Yang

Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

HierVST: Hierarchical Adaptive Zero-shot Voice Style Transfer

2023-07-30 · Sang-Hoon Lee, Ha-Yeong Choi, Hyung-Seok Oh, Seong-Whan Lee

Despite rapid progress in the voice style transfer (VST) field, recent zero-shot VST systems still lack the ability to transfer the voice style of a novel speaker. In this paper, we present HierVST, a hierarchical adapti…

Style TransferVariational Inference