Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism
This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal modeling is attractive. However, due to the difficulty of the temporal modeling of singing voices, many recent SVS systems with an encoder-decoder-based model still rely on explicitly on duration information generated by additional modules. Although some studies perform simultaneous modeling using seq2seq models with an attention mechanism, they have insufficient robustness against temporal modeling. The proposed attention mechanism is designed to estimate the attention weights by considering the rhythm given by the musical score. Furthermore, several techniques are also introduced to improve the modeling performance of the singing voice. Experimental results indicated that the proposed model is effective in terms of both naturalness and robustness of timing.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderPositionRhythmSinging Voice SynthesisMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Synthesising Expressiveness in Peking Opera via Duration Informed Attention Network
This paper presents a method that generates expressive singing voice of Peking opera. The synthesis of expressive opera singing usually requires pitch contours to be extracted as the training data, which relies on techni…
XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System
This paper presents XiaoiceSing, a high-quality singing voice synthesis system which employs an integrated network for spectrum, F0 and duration modeling. We follow the main architecture of FastSpeech while proposing som…
RhythmSinging Voice SynthesisVocal Bursts Intensity PredictionSinging voice synthesis based on convolutional neural networks
The present paper describes a singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the…
Singing Voice SynthesisVocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musi…
VISinger: Variational Inference with Adversarial Learning for End-to-End Singing Voice Synthesis
In this paper, we propose VISinger, a complete end-to-end high-quality singing voice synthesis (SVS) system that directly generates audio waveform from lyrics and musical score. Our approach is inspired by VITS, which ad…
DecoderRhythmSinging Voice SynthesisVariational Inference