paper-with-me

홈 › Papers

Sequence-to-sequence Singing Voice Synthesis with Perceptual Entropy Loss

2020-10-22 · Jiatong Shi, Shuai Guo, Nan Huo, Yuekai Zhang, Qin Jin

The neural network (NN) based singing voice synthesis (SVS) systems require sufficient data to train well and are prone to over-fitting due to data scarcity. However, we often encounter data limitation problem in building SVS systems because of high data acquisition and annotation costs. In this work, we propose a Perceptual Entropy (PE) loss derived from a psycho-acoustic hearing model to regularize the network. With a one-hour open-source singing voice database, we explore the impact of the PE loss on various mainstream sequence-to-sequence models, including the RNN-based, transformer-based, and conformer-based models. Our experiments show that the PE loss can mitigate the over-fitting problem and significantly improve the synthesized singing quality reflected in objective and subjective evaluations.

📄 PDF Abstract BibTeX arXiv:2010.12024

Code (1)

SJTMusicTeam/SVS_system 공식 구현 pytorch

Tasks

Singing Voice Synthesis

Similar Papers 제목 키워드 기반

Singing voice synthesis based on convolutional neural networks

2019-04-15 · Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku 외

The present paper describes a singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the…

Singing Voice Synthesis

Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence

2025-09-08 · Yerin Ryu, Inseop Shin, Chanwoo Kim arxiv

Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise con…

Fast and High-Quality Singing Voice Synthesis System based on Convolutional Neural Networks

2019-10-24 · Kazuhiro Nakamura, Shinji Takaki, Kei Hashimoto, Keiichiro Oura 외

The present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the n…

Singing Voice Synthesis

Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism

2022-12-28 · Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda

This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acous…

DecoderPositionRhythmSinging Voice Synthesis

HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis

2020-09-03 · Jiawei Chen, Xu Tan, Jian Luan, Tao Qin 외

High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and longer waveform sequences and throws cha…

Singing Voice SynthesisVocal Bursts Intensity Prediction