paper-with-me

Papers

Source-Filter HiFi-GAN: Fast and Pitch Controllable High-Fidelity Neural Vocoder

2022-10-27 · Reo Yoneyama, Yi-Chiao Wu, Tomoki Toda

Our previous work, the unified source-filter GAN (uSFGAN) vocoder, introduced a novel architecture based on the source-filter theory into the parallel waveform generative adversarial network to achieve high voice quality and pitch controllability. However, the high temporal resolution inputs result in high computation costs. Although the HiFi-GAN vocoder achieves fast high-fidelity voice generation thanks to the efficient upsampling-based generator architecture, the pitch controllability is severely limited. To realize a fast and pitch-controllable high-fidelity neural vocoder, we introduce the source-filter theory into HiFi-GAN by hierarchically conditioning the resonance filtering network on a well-estimated source excitation information. According to the experimental results, our proposed method outperforms HiFi-GAN and uSFGAN on a singing voice generation in voice quality and synthesis speed on a single CPU. Furthermore, unlike the uSFGAN vocoder, the proposed method can be easily adopted/integrated in real-time applications and end-to-end systems.

📄 PDF Abstract BibTeX arXiv:2210.15533

Code (0)

등록된 구현이 없습니다.

Tasks

CPUGenerative Adversarial Network

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
HiFi-GAN HiFi-GAN is a generative adversarial network for speech synthesis. HiFi-GAN consists of one generator and two discriminators: multi-scale and multi-period discriminators. The…

Similar Papers 제목 키워드 기반

Enhancement of Pitch Controllability using Timbre-Preserving Pitch Augmentation in FastPitch

2022-04-12 · Hanbin Bae, Young-Sun Joo

The recently developed pitch-controllable text-to-speech (TTS) model, i.e. FastPitch, was conditioned for the pitch contours. However, the quality of the synthesized speech degraded considerably for pitch values that dev…

Sentencetext-to-speechText to Speech

FastPitchFormant: Source-filter based Decomposed Modeling for Speech Synthesis

2021-06-29 · Taejun Bak, Jae-Sung Bae, Hanbin Bae, Young-Ik Kim 외

Methods for modeling and controlling prosody with acoustic features have been proposed for neural text-to-speech (TTS) models. Prosodic speech can be generated by conditioning acoustic features. However, synthesized spee…

Speech Synthesistext-to-speechText to Speech

HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform

2023-09-18 · Yinghao Aaron Li, Cong Han, Xilin Jiang, Nima Mesgarani

Recent advancements in speech synthesis have leveraged GAN-based networks like HiFi-GAN and BigVGAN to produce high-fidelity waveforms from mel-spectrograms. However, these networks are computationally expensive and para…

Speech Synthesis

HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis

2020-09-03 · Jiawei Chen, Xu Tan, Jian Luan, Tao Qin 외

High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and longer waveform sequences and throws cha…

Singing Voice SynthesisVocal Bursts Intensity Prediction

HiFi-Stream: Streaming Speech Enhancement with Generative Adversarial Networks

2025-03-21 · Ekaterina Dmitrieva, Maksim Kaledin

Speech Enhancement techniques have become core technologies in mobile devices and voice software simplifying downstream speech tasks. Still, modern Deep Learning (DL) solutions often require high amount of computational …

Speech Enhancement