paper-with-me

Papers

Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters

2024-01-10 · Kenichi Fujita, Hiroshi Sato, Takanori Ashihara, Hiroki Kanagawa, Marc Delcroix, Takafumi Moriya, Yusuke Ijima

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this approach suffers from degradation in speech synthesis quality when the reference speech contains noise. In this paper, we propose a noise-robust zero-shot TTS method. We incorporated adapters into the SSL model, which we fine-tuned with the TTS model using noisy reference speech. In addition, to further improve performance, we adopted a speech enhancement (SE) front-end. With these improvements, our proposed SSL-based zero-shot TTS achieved high-quality speech synthesis with noisy reference speech. Through the objective and subjective evaluations, we confirmed that the proposed method is highly robust to noise in reference speech, and effectively works in combination with SE.

📄 PDF Abstract BibTeX arXiv:2401.05111

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSpeech EnhancementSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron

2025-01-10 · Kishor Kayyar Lakshminarayana, Frank Zalkow, Christian Dittmar, Nicola Pia 외

In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many di…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

2023-11-21 · Sang-Hoon Lee, Ha-Yeong Choi, Seung-bin Kim, Seong-Whan Lee

Large language models (LLM)-based speech synthesis has been widely adopted in zero-shot speech synthesis. However, they require a large-scale data and possess the same limitations as previous autoregressive speech models…

Speech SynthesisSuper-Resolutiontext-to-speechText to Speech+2

StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis

2024-09-24 · Zhiyong Chen, Xinnuo Li, Zhiqi Ai, Shugong Xu

We introduce StyleFusion-TTS, a prompt and/or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research lite…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

FlashSpeech: Efficient Zero-Shot Speech Synthesis

2024-04-23 · Zhen Ye, Zeqian Ju, Haohe Liu, Xu Tan 외

Recent progress in large-scale zero-shot speech synthesis has been significantly advanced by language models and diffusion models. However, the generation process of both methods is slow and computationally intensive. Ef…

RhythmSpeech SynthesisVoice Conversion

NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

2023-04-18 · Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu 외

Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, and styles (e.g., singing). Current large …

In-Context LearningSpeech Synthesistext-to-speechText to Speech