paper-with-me

Papers

A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions

2024-11-19 · Hui-Peng Du, Ye-Xin Lu, Yang Ai, Zhen-Hua Ling

This paper proposes a novel neural denoising vocoder that can generate clean speech waveforms from noisy mel-spectrograms. The proposed neural denoising vocoder consists of two components, i.e., a spectrum predictor and a enhancement module. The spectrum predictor first predicts the noisy amplitude and phase spectra from the input noisy mel-spectrogram, and subsequently the enhancement module recovers the clean amplitude and phase spectrum from noisy ones. Finally, clean speech waveforms are reconstructed through inverse short-time Fourier transform (iSTFT). All operations are performed at the frame-level spectral domain, with the APNet vocoder and MP-SENet speech enhancement model used as the backbones for the two components, respectively. Experimental results demonstrate that our proposed neural denoising vocoder achieves state-of-the-art performance compared to existing neural vocoders on the VoiceBank+DEMAND dataset. Additionally, despite the lack of phase information and partial amplitude information in the input mel-spectrogram, the proposed neural denoising vocoder still achieves comparable performance with the serveral advanced speech enhancement methods.

📄 PDF Abstract BibTeX arXiv:2411.12268

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeech Enhancement

Similar Papers 제목 키워드 기반

Parametric Resynthesis with neural vocoders

2019-06-16 · Soumi Maiti, Michael I Mandel

Noise suppression systems generally produce output speech with compromised quality. We propose to utilize the high quality speech generation capability of neural vocoders for noise suppression. We use a neural network to…

Resynthesis

PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model

2024-02-22 · Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda

This paper presents a neural vocoder based on a denoising diffusion probabilistic model (DDPM) incorporating explicit periodic signals as auxiliary conditioning signals. Recently, DDPM-based neural vocoders have gained p…

DenoisingPitch controlSinging Voice Synthesis

CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR

2025-02-27 · Nian Shao, Rui Zhou, Pengyu Wang, Xian Li 외

In this work, we propose CleanMel, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) performance. The proposed network takes a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingSpeech Enhancement+2

LA-VocE: Low-SNR Audio-visual Speech Enhancement using Neural Vocoders

2022-11-20 · Rodrigo Mira, Buye Xu, Jacob Donley, Anurag Kumar 외

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvement…

Speech EnhancementSpeech Synthesis

BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation

2024-06-04 · Hui-Peng Du, Ye-Xin Lu, Yang Ai, Zhen-Hua Ling

This paper proposes a novel bidirectional neural vocoder, named BiVocoder, capable both of feature extraction and reverse waveform generation within the short-time Fourier transform (STFT) domain. For feature extraction,…

text-to-speechText to Speech