paper-with-me

Papers

Source-Filter-Based Generative Adversarial Neural Vocoder for High Fidelity Speech Synthesis

2023-04-26 · Ye-Xin Lu, Yang Ai, Zhen-Hua Ling

This paper proposes a source-filter-based generative adversarial neural vocoder named SF-GAN, which achieves high-fidelity waveform generation from input acoustic features by introducing F0-based source excitation signals to a neural filter framework. The SF-GAN vocoder is composed of a source module and a resolution-wise conditional filter module and is trained based on generative adversarial strategies. The source module produces an excitation signal from the F0 information, then the resolution-wise convolutional filter module combines the excitation signal with processed acoustic features at various temporal resolutions and finally reconstructs the raw waveform. The experimental results show that our proposed SF-GAN vocoder outperforms the state-of-the-art HiFi-GAN and Fre-GAN in both analysis-synthesis (AS) and text-to-speech (TTS) tasks, and the synthesized speech quality of SF-GAN is comparable to the ground-truth audio.

📄 PDF Abstract BibTeX arXiv:2304.13270

Code (1)

yxlu-0102/MP-SENet pytorch

Tasks

Speech Synthesistext-to-speechText to Speech

Methods 이 논문이 사용한 방법론

HiFi-GAN HiFi-GAN is a generative adversarial network for speech synthesis. HiFi-GAN consists of one generator and two discriminators: multi-scale and multi-period discriminators. The…

Similar Papers 제목 키워드 기반

Source-Filter HiFi-GAN: Fast and Pitch Controllable High-Fidelity Neural Vocoder

2022-10-27 · Reo Yoneyama, Yi-Chiao Wu, Tomoki Toda

Our previous work, the unified source-filter GAN (uSFGAN) vocoder, introduced a novel architecture based on the source-filter theory into the parallel waveform generative adversarial network to achieve high voice quality…

CPUGenerative Adversarial Network

Unified Source-Filter GAN: Unified Source-filter Network Based On Factorization of Quasi-Periodic Parallel WaveGAN

2021-04-10 · Reo Yoneyama, Yi-Chiao Wu, Tomoki Toda

We propose a unified approach to data-driven source-filter modeling using a single neural network for developing a neural vocoder capable of generating high-quality synthetic speech waveforms while retaining flexibility …

A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis

2018-04-07 · Xin Wang, Jaime Lorenzo-Trueba, Shinji Takaki, Lauri Juvela 외

Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using…

Speech Synthesis

Avocodo: Generative Adversarial Network for Artifact-free Vocoder

2022-06-27 · Taejun Bak, Junmo Lee, Hanbin Bae, Jinhyeok Yang 외

Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptu…

Generative Adversarial Network

WaveFit: An Iterative and Non-autoregressive Neural Vocoder based on Fixed-Point Iteration

2022-10-03 · Yuma Koizumi, Kohei Yatabe, Heiga Zen, Michiel Bacchiani

Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characterized by the iterative denoising framework …

Denoising