paper-with-me

Papers

Peking Opera Synthesis via Duration Informed Attention Network

2020-08-07 · Yusong Wu, Shengchen Li, Chengzhu Yu, Heng Lu, Chao Weng, Liqiang Zhang, Dong Yu

Peking Opera has been the most dominant form of Chinese performing art since around 200 years ago. A Peking Opera singer usually exhibits a very strong personal style via introducing improvisation and expressiveness on stage which leads the actual rhythm and pitch contour to deviate significantly from the original music score. This inconsistency poses a great challenge in Peking Opera singing voice synthesis from a music score. In this work, we propose to deal with this issue and synthesize expressive Peking Opera singing from the music score based on the Duration Informed Attention Network (DurIAN) framework. To tackle the rhythm mismatch, Lagrange multiplier is used to find the optimal output phoneme duration sequence with the constraint of the given note duration from music score. As for the pitch contour mismatch, instead of directly inferring from music score, we adopt a pseudo music score generated from the real singing and feed it as input during training. The experiments demonstrate that with the proposed system we can synthesize Peking Opera singing voice with high-quality timbre, pitch and expressiveness.

📄 PDF Abstract BibTeX arXiv:2008.03029

Code (0)

등록된 구현이 없습니다.

Tasks

RhythmSinging Voice Synthesis

Similar Papers 제목 키워드 기반

Synthesising Expressiveness in Peking Opera via Duration Informed Attention Network

2019-12-27 · Yusong Wu, Shengchen Li, Chengzhu Yu, Heng Lu 외

This paper presents a method that generates expressive singing voice of Peking opera. The synthesis of expressive opera singing usually requires pitch contours to be extracted as the training data, which relies on techni…

DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis

2023-09-22 · Yu Gu, Yianrao Bian, Guangzhi Lei, Chao Weng 외

This paper introduces an improved duration informed attention neural network (DurIAN-E) for expressive and high-fidelity text-to-speech (TTS) synthesis. Inherited from the original DurIAN model, an auto-regressive model …

DenoisingSpeech Synthesistext-to-speechText to Speech+1

DurIAN: Duration Informed Attention Network For Multimodal Synthesis

2019-09-04 · Chengzhu Yu, Heng Lu, Na Hu, Meng Yu 외

In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously. The key component of this system is the Duration Informed Attention Net…

CPUSpeech Synthesis

DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis

2024-10-17 · Yu Gu, Qiushi Zhu, Guangzhi Lei, Chao Weng 외

This paper proposes an improved version of DurIAN-E (DurIAN-E 2), which is also a duration informed attention neural network for expressive and high-fidelity text-to-speech (TTS) synthesis. Similar with the DurIAN-E mode…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

PAMA-TTS: Progression-Aware Monotonic Attention for Stable Seq2Seq TTS With Accurate Phoneme Duration Control

2021-10-09 · Yunchao He, Jian Luan, Yujun Wang

Sequence expansion between encoder and decoder is a critical challenge in sequence-to-sequence TTS. Attention-based methods achieve great naturalness but suffer from unstable issues like missing and repeating phonemes, n…

Decoder