paper-with-me

홈 › Papers

Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism

2022-12-28 · Yukiya Hono, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda

This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal modeling is attractive. However, due to the difficulty of the temporal modeling of singing voices, many recent SVS systems with an encoder-decoder-based model still rely on explicitly on duration information generated by additional modules. Although some studies perform simultaneous modeling using seq2seq models with an attention mechanism, they have insufficient robustness against temporal modeling. The proposed attention mechanism is designed to estimate the attention weights by considering the rhythm given by the musical score. Furthermore, several techniques are also introduced to improve the modeling performance of the singing voice. Experimental results indicated that the proposed model is effective in terms of both naturalness and robustness of timing.

📄 PDF Abstract BibTeX arXiv:2212.13703

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderPositionRhythmSinging Voice Synthesis

Methods 이 논문이 사용한 방법론

Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…
Seq2Seq Seq2Seq, or Sequence To Sequence, is a model used in sequence prediction tasks, such as language modelling and machine translation. The idea is to use one…

Similar Papers 제목 키워드 기반

Synthesising Expressiveness in Peking Opera via Duration Informed Attention Network

2019-12-27 · Yusong Wu, Shengchen Li, Chengzhu Yu, Heng Lu 외

This paper presents a method that generates expressive singing voice of Peking opera. The synthesis of expressive opera singing usually requires pitch contours to be extracted as the training data, which relies on techni…

XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System

2020-06-11 · Peiling Lu, Jie Wu, Jian Luan, Xu Tan 외

This paper presents XiaoiceSing, a high-quality singing voice synthesis system which employs an integrated network for spectrum, F0 and duration modeling. We follow the main architecture of FastSpeech while proposing som…

RhythmSinging Voice SynthesisVocal Bursts Intensity Prediction

Singing voice synthesis based on convolutional neural networks

2019-04-15 · Kazuhiro Nakamura, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku 외

The present paper describes a singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currently being proposed and are improving the…

Singing Voice Synthesis

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models

2026-05-06 · Yukun Chen, Tianrui Wang, Zhaoxi Mu, Xinyu Yang 외 arxiv

High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musi…

VISinger: Variational Inference with Adversarial Learning for End-to-End Singing Voice Synthesis

2021-10-17 · Yongmao Zhang, Jian Cong, Heyang Xue, Lei Xie 외

In this paper, we propose VISinger, a complete end-to-end high-quality singing voice synthesis (SVS) system that directly generates audio waveform from lyrics and musical score. Our approach is inspired by VITS, which ad…

DecoderRhythmSinging Voice SynthesisVariational Inference