paper-with-me

Papers

Avocodo: Generative Adversarial Network for Artifact-free Vocoder

2022-06-27 · Taejun Bak, Junmo Lee, Hanbin Bae, Jinhyeok Yang, Jae-Sung Bae, Young-Sun Joo

Neural vocoders based on the generative adversarial neural network (GAN) have been widely used due to their fast inference speed and lightweight networks while generating high-quality speech waveforms. Since the perceptually important speech components are primarily concentrated in the low-frequency bands, most GAN-based vocoders perform multi-scale analysis that evaluates downsampled speech waveforms. This multi-scale analysis helps the generator improve speech intelligibility. However, in preliminary experiments, we discovered that the multi-scale analysis which focuses on the low-frequency bands causes unintended artifacts, e.g., aliasing and imaging artifacts, which degrade the synthesized speech waveform quality. Therefore, in this paper, we investigate the relationship between these artifacts and GAN-based vocoders and propose a GAN-based vocoder, called Avocodo, that allows the synthesis of high-fidelity speech with reduced artifacts. We introduce two kinds of discriminators to evaluate speech waveforms in various perspectives: a collaborative multi-band discriminator and a sub-band discriminator. We also utilize a pseudo quadrature mirror filter bank to obtain downsampled multi-band speech waveforms while avoiding aliasing. According to experimental results, Avocodo outperforms baseline GAN-based vocoders, both objectively and subjectively, while reproducing speech with fewer artifacts.

📄 PDF Abstract BibTeX arXiv:2206.13404

Code (2)

ncsoft/Avocodo 공식 구현 pytorch
rishikksh20/Avocodo-pytorch pytorch

Tasks

Generative Adversarial Network

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

FA-GAN: Artifacts-free and Phase-aware High-fidelity GAN-based Vocoder

2024-07-05 · Rubing Shen, Yanzhen Ren, Zongkun Sun

Generative adversarial network (GAN) based vocoders have achieved significant attention in speech synthesis with high quality and fast inference speed. However, there still exist many noticeable spectral artifacts, resul…

Generative Adversarial NetworkSpeech Synthesis

Source-Filter-Based Generative Adversarial Neural Vocoder for High Fidelity Speech Synthesis

2023-04-26 · Ye-Xin Lu, Yang Ai, Zhen-Hua Ling

This paper proposes a source-filter-based generative adversarial neural vocoder named SF-GAN, which achieves high-fidelity waveform generation from input acoustic features by introducing F0-based source excitation signal…

Speech Synthesistext-to-speechText to Speech

A Post Auto-regressive GAN Vocoder Focused on Spectrum Fracture

2022-04-12 · Zhenxing Lu, Mengnan He, Ruixiong Zhang, Caixia Gong

Generative adversarial networks (GANs) have been indicated their superiority in usage of the real-time speech synthesis. Nevertheless, most of them make use of deep convolutional layers as their backbone, which may cause…

Speech Synthesis

Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis

2022-06-23 · Tae-Woo Kim, Min-Su Kang, Gyeong-Hoon Lee

Recently, deep learning-based generative models have been introduced to generate singing voices. One approach is to predict the parametric vocoder features consisting of explicit speech parameters. This approach has the …

Generative Adversarial NetworkMulti-Task LearningSinging Voice Synthesis

Towards generalizing deep-audio fake detection networks

2023-05-22 · Konstantin Gasenzer, Moritz Wolter

Today's generative neural networks allow the creation of high-quality synthetic speech at scale. While we welcome the creative use of this new technology, we must also recognize the risks. As synthetic speech is abused f…

Face Swapping