paper-with-me

Papers

Speech waveform synthesis from MFCC sequences with generative adversarial networks

2018-04-03 · Lauri Juvela, Bajibabu Bollepalli, Xin Wang, Hirokazu Kameoka, Manu Airaksinen, Junichi Yamagishi, Paavo Alku

This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from MFCCs with an autoregressive recurrent neural net. Second, the spectral envelope information contained in MFCCs is converted to all-pole filters, and a pitch-synchronous excitation model matched to these filters is trained. Finally, we introduce a generative adversarial network -based noise model to add a realistic high-frequency stochastic component to the modeled excitation signal. The results show that high quality speech reconstruction can be obtained, given only MFCC information at test time.

📄 PDF Abstract BibTeX arXiv:1804.00920

Code (1)

ljuvela/ResGAN 공식 구현 tf

Tasks

Generative Adversarial NetworkSpeech Synthesis

Similar Papers 제목 키워드 기반

WARP-Q: Quality Prediction For Generative Neural Speech Codecs

2021-02-20 · Wissam A. Jassim, Jan Skoglund, Michael Chinen, Andrew Hines

Good speech quality has been achieved using waveform matching and parametric reconstruction coders. Recently developed very low bit rate generative codecs can reconstruct high quality wideband speech with bit streams les…

Dynamic Time WarpingPrediction

Speech Recognition Front End Without Information Loss

2013-12-24 · Matthew Ager, Zoran Cvetkovic, Peter Sollich

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)General Classificationspeech-recognition+1

MFCCGAN: A Novel MFCC-Based Speech Synthesizer Using Adversarial Learning

2023-06-22 · Mohammad Reza Hasanabadi Majid Behdad Davood Gharavian

In this paper, we introduce MFCCGAN as a novel speech synthesizer based on adversarial learning that adopts MFCCs as input and generates raw speech waveforms. Benefiting the GAN model capabilities, it produces speech wit…

Multi-View Speech Representation Learning for Parkinson's Disease Detection Using Context-guided Cross-modal Attention

2026-06-08 · George Theodosiou, Loukas Ilias, Dimitris Askounis arxiv

Parkinson's disease (PD) is a progressive neurodegenerative disorder that frequently causes speech impairments associated with hypokinetic dysarthria. As speech production relies on the precise coordination of complex ne…

Representation Learning

Generative adversarial network-based glottal waveform model for statistical parametric speech synthesis

2019-03-14 · Bajibabu Bollepalli, Lauri Juvela, Paavo Alku

Recent studies have shown that text-to-speech synthesis quality can be improved by using glottal vocoding. This refers to vocoders that parameterize speech into two parts, the glottal excitation and vocal tract, that occ…

Generative Adversarial NetworkSpeech Synthesistext-to-speechText to Speech+1