paper-with-me

홈 › Papers

JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech

2022-03-31 · Dan Lim, Sunghee Jung, Eesung Kim

In neural text-to-speech (TTS), two-stage system or a cascade of separately learned models have shown synthesis quality close to human speech. For example, FastSpeech2 transforms an input text to a mel-spectrogram and then HiFi-GAN generates a raw waveform from a mel-spectogram where they are called an acoustic feature generator and a neural vocoder respectively. However, their training pipeline is somewhat cumbersome in that it requires a fine-tuning and an accurate speech-text alignment for optimal performance. In this work, we present end-to-end text-to-speech (E2E-TTS) model which has a simplified training pipeline and outperforms a cascade of separately learned models. Specifically, our proposed model is jointly trained FastSpeech2 and HiFi-GAN with an alignment module. Since there is no acoustic feature mismatch between training and inference, it does not requires fine-tuning. Furthermore, we remove dependency on an external speech-text alignment tool by adopting an alignment learning objective in our joint training framework. Experiments on LJSpeech corpus shows that the proposed model outperforms publicly available, state-of-the-art implementations of ESPNet2-TTS on subjective evaluation (MOS) and some objective evaluations.

📄 PDF Abstract BibTeX arXiv:2203.16852

Code (2)

imdanboy/jets 공식 구현 pytorch
keonlee9420/Comprehensive-E2E-TTS pytorch

Tasks

text-to-speechText to Speech

Methods 이 논문이 사용한 방법론

HiFi-GAN HiFi-GAN is a generative adversarial network for speech synthesis. HiFi-GAN consists of one generator and two discriminators: multi-scale and multi-period discriminators. The…

Similar Papers 제목 키워드 기반

HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis

2020-09-03 · Jiawei Chen, Xu Tan, Jian Luan, Tao Qin 외

High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and longer waveform sequences and throws cha…

Singing Voice SynthesisVocal Bursts Intensity Prediction

PRESENT: Zero-Shot Text-to-Prosody Control

2024-08-13 · Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman 외

Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-t…

Prosody PredictionSpeech Synthesistext-to-speechText to Speech

FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

2020-06-08 · ICLR 2021 1 · Yi Ren, Chenxu Hu, Xu Tan, Tao Qin 외

Non-autoregressive text to speech (TTS) models such as FastSpeech can synthesize speech significantly faster than previous autoregressive models with comparable quality. The training of FastSpeech model relies on an auto…

Knowledge DistillationSpeech Synthesistext-to-speechText to Speech+1

MnTTS2: An Open-Source Multi-Speaker Mongolian Text-to-Speech Synthesis Dataset

2022-12-11 · Kailin Liang, Bin Liu, Yifan Hu, Rui Liu 외

Text-to-Speech (TTS) synthesis for low-resource languages is an attractive research issue in academia and industry nowadays. Mongolian is the official language of the Inner Mongolia Autonomous Region and a representative…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

MnTTS: An Open-Source Mongolian Text-to-Speech Synthesis Dataset and Accompanied Baseline

2022-09-22 · Yifan Hu, Pengkai Yin, Rui Liu, Feilong Bao 외

This paper introduces a high-quality open-source text-to-speech (TTS) synthesis dataset for Mongolian, a low-resource language spoken by over 10 million people worldwide. The dataset, named MnTTS, consists of about 8 hou…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis