paper-with-me

홈 › Papers

PeriodNet: A non-autoregressive waveform generation model with a structure separating periodic and aperiodic components

2021-02-15 · Yukiya Hono, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, Keiichi Tokuda

We propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non-AR waveform generation models can generate speech waveforms parallelly and can be used as a speech vocoder by conditioning an acoustic feature. Since a speech waveform contains periodic and aperiodic components, both components should be appropriately modeled to generate a high-quality speech waveform. However, it is difficult to decompose the components from a natural speech waveform in advance. To address this issue, we propose a parallel model and a series model structure separating periodic and aperiodic components. The features of our proposed models are that explicit periodic and aperiodic signals are taken as input, and external periodic/aperiodic decomposition is not needed in training. Experiments using a singing voice corpus show that our proposed structure improves the naturalness of the generated waveform. We also show that the speech waveforms with a pitch outside of the training data range can be generated with more naturalness.

📄 PDF Abstract BibTeX arXiv:2102.07786

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PeriodNet: Boosting the Potential of Attention Mechanism for Time Series Forecasting

2025-11-23 · Bowen Zhao, Huanlai Xing, Zhiwen Xiao, Jincheng Peng 외 arxiv

The attention mechanism has demonstrated remarkable potential in sequence modeling, exemplified by its successful application in natural language processing with models such as Bidirectional Encoder Representations from …

Time Series Forecasting

DiffWave: A Versatile Diffusion Model for Audio Synthesis

2020-09-21 · ICLR 2021 1 · Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 외

In this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation. The model is non-autoregressive, and converts the white noise signal into structured wav…

Audio SynthesisDiversitymodelSpeech Synthesis

It's Raw! Audio Generation with State-Space Models

2022-02-20 · Karan Goel, Albert Gu, Chris Donahue, Christopher Ré

Developing architectures suitable for modeling raw audio is a challenging problem due to the high sampling rates of audio waveforms. Standard sequence modeling approaches like RNNs and CNNs have previously been tailored …

Audio GenerationDensity EstimationMusic GenerationState Space Models

The challenge of realistic music generation: modelling raw audio at scale

2018-06-26 · NeurIPS 2018 12 · Sander Dieleman, Aäron van den Oord, Karen Simonyan

Realistic music generation is a challenging task. When building generative models of music that are learnt from data, typically high-level representations such as scores or MIDI are used that abstract away the idiosyncra…

Music Generation

Sinsy: A Deep Neural Network-Based Singing Voice Synthesis System

2021-08-05 · Yukiya Hono, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku 외

This paper presents Sinsy, a deep neural network (DNN)-based singing voice synthesis (SVS) system. In recent years, DNNs have been utilized in statistical parametric SVS systems, and DNN-based SVS systems have demonstrat…

Singing Voice Synthesis