paper-with-me

홈 › Papers

Singing voice synthesis based on frame-level sequence-to-sequence models considering vocal timing deviation

2023-01-05 · Miku Nishihara, Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda

This paper proposes singing voice synthesis (SVS) based on frame-level sequence-to-sequence models considering vocal timing deviation. In SVS, it is essential to synchronize the timing of singing with temporal structures represented by scores, taking into account that there are differences between actual vocal timing and note start timing. In many SVS systems including our previous work, phoneme-level score features are converted into frame-level ones on the basis of phoneme boundaries obtained by external aligners to take into account vocal timing deviations. Therefore, the sound quality is affected by the aligner accuracy in this system. To alleviate this problem, we introduce an attention mechanism with frame-level features. In the proposed system, the attention mechanism absorbs alignment errors in phoneme boundaries. Additionally, we evaluate the system with pseudo-phoneme-boundaries defined by heuristic rules based on musical scores when there is no aligner. The experimental results show the effectiveness of the proposed system.

📄 PDF Abstract BibTeX arXiv:2301.02262

Code (0)

등록된 구현이 없습니다.

Tasks

Singing Voice Synthesis

Similar Papers 제목 키워드 기반

Singing voice synthesis based on convolutional neural networks

2019-04-15 · Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku 외

The present paper describes a singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the…

Singing Voice Synthesis

Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence

2025-09-08 · Yerin Ryu, Inseop Shin, Chanwoo Kim arxiv

Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise con…

Fast and High-Quality Singing Voice Synthesis System based on Convolutional Neural Networks

2019-10-24 · Kazuhiro Nakamura, Shinji Takaki, Kei Hashimoto, Keiichiro Oura 외

The present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the n…

Singing Voice Synthesis

Robust Singing Voice Transcription Serves Synthesis

2024-05-16 · RuiQi Li, Yu Zhang, Yongqi Wang, Zhiqing Hong 외

Note-level Automatic Singing Voice Transcription (AST) converts singing recordings into note sequences, facilitating the automatic annotation of singing datasets for Singing Voice Synthesis (SVS) applications. Current AS…

DecoderSinging Voice Synthesis

Sequence-to-sequence Singing Voice Synthesis with Perceptual Entropy Loss

2020-10-22 · Jiatong Shi, Shuai Guo, Nan Huo, Yuekai Zhang 외

The neural network (NN) based singing voice synthesis (SVS) systems require sufficient data to train well and are prone to over-fitting due to data scarcity. However, we often encounter data limitation problem in buildin…

Singing Voice Synthesis