paper-with-me

Papers

Disentangling Correlated Speaker and Noise for Speech Synthesis via Data Augmentation and Adversarial Factorization

2018-10-30 · NIPS Workshop IRASL 2018 · Anonymous

To leverage crowd-sourced data to train multi-speaker text-to-speech (TTS) models that can synthesize clean speech for all speakers, it is essential to learn disentangled representations which can independently control the speaker identity and background noise in generated signals. However, learning such representations can be challenging, due to the lack of labels describing the recording conditions of each training example, and the fact that speakers and recording conditions are often correlated, e.g. since users often make many recordings using the same equipment. This paper proposes three components to address this problem by: (1) formulating a conditional generative model with factorized latent variables, (2) using data augmentation to add noise that is not correlated with speaker identity and whose label is known during training, and (3) using adversarial factorization to improve disentanglement. Experimental results demonstrate that the proposed method can disentangle speaker and noise attributes even if they are correlated in the training data, and can be used to consistently synthesize clean speech for all speakers. Ablation studies verify the importance of each proposed component.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDisentanglementSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement

2023-06-27 · Siqi Zheng, Luyao Cheng, Yafeng Chen, Hui Wang 외

Disentangling uncorrelated information in speech utterances is a crucial research topic within speech community. Different speech-related tasks focus on extracting distinct speech representations while minimizing the aff…

DisentanglementSelf-Supervised Learning

Noise Robust TTS for Low Resource Speakers using Pre-trained Model and Speech Enhancement

2020-05-26 · Dongyang Dai, Li Chen, Yu-Ping Wang, Mu Wang 외

With the popularity of deep neural network, speech synthesis task has achieved significant improvements based on the end-to-end encoder-decoder framework in the recent days. More and more applications relying on speech s…

DecoderSpeech EnhancementSpeech Synthesis

CrossSpeech: Speaker-independent Acoustic Representation for Cross-lingual Speech Synthesis

2023-02-28 · Ji-Hoon Kim, Hong-Sun Yang, Yoon-Cheol Ju, Il-Hwan Kim 외

While recent text-to-speech (TTS) systems have made remarkable strides toward human-level quality, the performance of cross-lingual TTS lags behind that of intra-lingual TTS. This gap is mainly rooted from the speaker-la…

Speech Synthesistext-to-speechText to Speech

SelfVC: Voice Conversion With Iterative Refinement using Self Transformations

2023-10-14 · Paarth Neekhara, Shehzeen Hussain, Rafael Valle, Boris Ginsburg 외

We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled represe…

Self-Supervised LearningSpeaker VerificationSpeech SynthesisVoice Conversion

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron

2025-01-10 · Kishor Kayyar Lakshminarayana, Frank Zalkow, Christian Dittmar, Nicola Pia 외

In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many di…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis