paper-with-me

Papers

Training a Neural Speech Waveform Model using Spectral Losses of Short-Time Fourier Transform and Continuous Wavelet Transform

2019-03-29 · Shinji Takaki, Hirokazu Kameoka, Junichi Yamagishi

Recently, we proposed short-time Fourier transform (STFT)-based loss functions for training a neural speech waveform model. In this paper, we generalize the above framework and propose a training scheme for such models based on spectral amplitude and phase losses obtained by either STFT or continuous wavelet transform (CWT), or both of them. Since CWT is capable of having time and frequency resolutions different from those of STFT and is cable of considering those closer to human auditory scales, the proposed loss functions could provide complementary information on speech signals. Experimental results showed that it is possible to train a high-quality model by using the proposed CWT spectral loss and is as good as one using STFT-based loss.

📄 PDF Abstract BibTeX arXiv:1903.12392

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

APNet2: High-quality and High-efficiency Neural Vocoder with Direct Prediction of Amplitude and Phase Spectra

2023-11-20 · Hui-Peng Du, Ye-Xin Lu, Yang Ai, Zhen-Hua Ling

In our previous work, we proposed a neural vocoder called APNet, which directly predicts speech amplitude and phase spectra with a 5 ms frame shift in parallel from the input acoustic features, and then reconstructs the …

Speech Synthesis

STFT spectral loss for training a neural speech waveform model

2018-10-29 · Shinji Takaki, Toru Nakashika, Xin Wang, Junichi Yamagishi

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not …

Speech Denoising with Auditory Models

2020-11-21 · Mark R. Saddler, Andrew Francl, Jenelle Feather, Kaizhi Qian 외

Contemporary speech enhancement predominantly relies on audio transforms that are trained to reconstruct a clean speech waveform. The development of high-performing neural network sound recognition systems has raised the…

DenoisingSpeech DenoisingSpeech Enhancement

A Conformer-based Waveform-domain Neural Acoustic Echo Canceller Optimized for ASR Accuracy

2022-05-06 · Sankaran Panchapagesan, Arun Narayanan, Turaj Zakizadeh Shabestary, Shuai Shao 외

Acoustic Echo Cancellation (AEC) is essential for accurate recognition of queries spoken to a smart speaker that is playing out audio. Previous work has shown that a neural AEC model operating on log-mel spectral feature…

Acoustic echo cancellationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1

Neural source-filter waveform models for statistical parametric speech synthesis

2019-04-27 · Xin Wang, Shinji Takaki, Junichi Yamagishi

Neural waveform models such as WaveNet have demonstrated better performance than conventional vocoders for statistical parametric speech synthesis. As an autoregressive (AR) model, WaveNet is limited by a slow sequential…

Speech Synthesis