paper-with-me

Papers

A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis

2024-06-18 · Guoqiang Hu, Huaning Tan, Ruilai Li

Acoustic features play an important role in improving the quality of the synthesised speech. Currently, the Mel spectrogram is a widely employed acoustic feature in most acoustic models. However, due to the fine-grained loss caused by its Fourier transform process, the clarity of speech synthesised by Mel spectrogram is compromised in mutant signals. In order to obtain a more detailed Mel spectrogram, we propose a Mel spectrogram enhancement paradigm based on the continuous wavelet transform (CWT). This paradigm introduces an additional task: a more detailed wavelet spectrogram, which like the post-processing network takes as input the Mel spectrogram output by the decoder. We choose Tacotron2 and Fastspeech2 for experimental validation in order to test autoregressive (AR) and non-autoregressive (NAR) speech systems, respectively. The experimental results demonstrate that the speech synthesised using the model with the Mel spectrogram enhancement paradigm exhibits higher MOS, with an improvement of 0.14 and 0.09 compared to the baseline model, respectively. These findings provide some validation for the universality of the enhancement paradigm, as they demonstrate the success of the paradigm in different architectures.

📄 PDF Abstract BibTeX arXiv:2406.12164

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSpeech Synthesis

Similar Papers 제목 키워드 기반

High-quality Speech Synthesis Using Super-resolution Mel-Spectrogram

2019-12-03 · Leyuan Sheng, Dong-Yan Huang, Evgeniy N. Pavlovskiy

In speech synthesis and speech enhancement systems, melspectrograms need to be precise in acoustic representations. However, the generated spectrograms are over-smooth, that could not produce high quality synthesized spe…

Image-to-Image TranslationSpeech EnhancementSpeech SynthesisSuper-Resolution+2

Diffusion-Based Mel-Spectrogram Enhancement for Personalized Speech Synthesis with Found Data

2023-05-18 · Yusheng Tian, Wei Liu, Tan Lee

Creating synthetic voices with found data is challenging, as real-world recordings often contain various types of audio degradation. One way to address this problem is to pre-enhance the speech with an enhancement model …

Speech EnhancementSpeech Synthesistext-to-speechText to Speech

Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR

2024-02-21 · Rui Zhou, Xian Li, Ying Fang, Xiaofei Li

In this work, we propose Mel-FullSubNet, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) performance. Mel-FullSubNet takes a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingSpeech Enhancement+2

End-to-End Model for Speech Enhancement by Consistent Spectrogram Masking

2019-01-02 · Xingjian Du, Mengyao Zhu, Xuan Shi, Xinpeng Zhang 외

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier tra…

Speech Enhancement

Complex spectrogram enhancement by convolutional neural network with multi-metrics learning

2017-04-27 · Szu-Wei Fu, Ting-yao Hu, Yu Tsao, Xugang Lu

This paper aims to address two issues existing in the current speech enhancement methods: 1) the difficulty of phase estimations; 2) a single objective function cannot consider multiple metrics simultaneously. To solve t…

Speech Enhancement