High-quality Speech Synthesis Using Super-resolution Mel-Spectrogram
In speech synthesis and speech enhancement systems, melspectrograms need to be precise in acoustic representations. However, the generated spectrograms are over-smooth, that could not produce high quality synthesized speech. Inspired by image-to-image translation, we address this problem by using a learning-based post filter combining Pix2PixHD and ResUnet to reconstruct the mel-spectrograms together with super-resolution. From the resulting super-resolution spectrogram networks, we can generate enhanced spectrograms to produce high quality synthesized speech. Our proposed model achieves improved mean opinion scores (MOS) of 3.71 and 4.01 over baseline results of 3.29 and 3.84, while using vocoder Griffin-Lim and WaveNet, respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
Image-to-Image TranslationSpeech EnhancementSpeech SynthesisSuper-ResolutionTranslationVocal Bursts Intensity PredictionSimilar Papers 제목 키워드 기반
HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis
Large language models (LLM)-based speech synthesis has been widely adopted in zero-shot speech synthesis. However, they require a large-scale data and possess the same limitations as previous autoregressive speech models…
Speech SynthesisSuper-Resolutiontext-to-speechText to Speech+2Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
Articulatory trajectories like electromagnetic articulography (EMA) provide a low-dimensional representation of the vocal tract filter and have been used as natural, grounded features for speech synthesis. Differentiable…
Audio SynthesisComputational EfficiencyCPUSpeech SynthesisEXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis
Recent work has shown that it is possible to resynthesize high-quality speech based, not on text, but on low bitrate discrete units that have been learned in a self-supervised fashion and can therefore capture expressive…
ResynthesisSpeech SynthesisSource-Filter-Based Generative Adversarial Neural Vocoder for High Fidelity Speech Synthesis
This paper proposes a source-filter-based generative adversarial neural vocoder named SF-GAN, which achieves high-fidelity waveform generation from input acoustic features by introducing F0-based source excitation signal…
Speech Synthesistext-to-speechText to SpeechSuper-resolution Using Constrained Deep Texture Synthesis
Hallucinating high frequency image details in single image super-resolution is a challenging task. Traditional super-resolution methods tend to produce oversmoothed output images due to the ambiguity in mapping between l…
Image Super-ResolutionSuper-ResolutionTexture Synthesis