paper-with-me

홈 › Papers

Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization

2024-08-15 · Sang-Hoon Lee, Ha-Yeong Choi, Seong-Whan Lee

This paper introduces PeriodWave-Turbo, a high-fidelity and high-efficient waveform generation model via adversarial flow matching optimization. Recently, conditional flow matching (CFM) generative models have been successfully adopted for waveform generation tasks, leveraging a single vector field estimation objective for training. Although these models can generate high-fidelity waveform signals, they require significantly more ODE steps compared to GAN-based models, which only need a single generation step. Additionally, the generated samples often lack high-frequency information due to noisy vector field estimation, which fails to ensure high-frequency reproduction. To address this limitation, we enhance pre-trained CFM-based generative models by incorporating a fixed-step generator modification. We utilized reconstruction losses and adversarial feedback to accelerate high-fidelity waveform generation. Through adversarial flow matching optimization, it only requires 1,000 steps of fine-tuning to achieve state-of-the-art performance across various objective metrics. Moreover, we significantly reduce inference speed from 16 steps to 2 or 4 steps. Additionally, by scaling up the backbone of PeriodWave from 29M to 70M parameters for improved generalization, PeriodWave-Turbo achieves unprecedented performance, with a perceptual evaluation of speech quality (PESQ) score of 4.454 on the LibriTTS dataset. Audio samples, source code and checkpoints will be available at https://github.com/sh-lee-prml/PeriodWave.

📄 PDF Abstract BibTeX arXiv:2408.08019

Code (1)

sh-lee-prml/periodwave 공식 구현 pytorch

Tasks

Speech Synthesis

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram

2019-10-25 · Ryuichi Yamamoto, Eunwoo Song, Jae-Min Kim

We propose Parallel WaveGAN, a distillation-free, fast, and small-footprint waveform generation method using a generative adversarial network. In the proposed method, a non-autoregressive WaveNet is trained by jointly op…

Generative Adversarial NetworkGPUSpeech Synthesistext-to-speech+2

Source-Filter-Based Generative Adversarial Neural Vocoder for High Fidelity Speech Synthesis

2023-04-26 · Ye-Xin Lu, Yang Ai, Zhen-Hua Ling

This paper proposes a source-filter-based generative adversarial neural vocoder named SF-GAN, which achieves high-fidelity waveform generation from input acoustic features by introducing F0-based source excitation signal…

Speech Synthesistext-to-speechText to Speech

TFGAN: Time and Frequency Domain Based Generative Adversarial Network for High-fidelity Speech Synthesis

2020-11-24 · Qiao Tian, Yi Chen, Zewang Zhang, Heng Lu 외

Recently, GAN based speech synthesis methods, such as MelGAN, have become very popular. Compared to conventional autoregressive based methods, parallel structures based generators make waveform generation process fast an…

Generative Adversarial NetworkSpeech Synthesis

RefineGAN: Universally Generating Waveform Better than Ground Truth with Highly Accurate Pitch and Intensity Responses

2021-11-01 · Shengyuan Xu, Wenxiao Zhao, Jing Guo

Most GAN(Generative Adversarial Network)-based approaches towards high-fidelity waveform generation heavily rely on discriminators to improve their performance. However, GAN methods introduce much uncertainty into the ge…

Audio GenerationGenerative Adversarial NetworkSinging Voice Synthesis

HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis

2020-09-03 · Jiawei Chen, Xu Tan, Jian Luan, Tao Qin 외

High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and longer waveform sequences and throws cha…

Singing Voice SynthesisVocal Bursts Intensity Prediction