paper-with-me

홈 › Papers

Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis

2022-03-02 · Pengyu Cheng, ZhenHua Ling

In this paper, we propose a method of speaker adaption with intuitive prosodic features for statistical parametric speech synthesis. The intuitive prosodic features employed in this method include pitch, pitch range, speech rate and energy considering that they are directly related with the overall prosodic characteristics of different speakers. The intuitive prosodic features are extracted at utterance-level or speaker-level, and are further integrated into the existing speaker-encoding-based and speaker-embedding-based adaptation frameworks respectively. The acoustic models are sequence-to-sequence ones based on Tacotron2. Intuitive prosodic features are concatenated with text encoder outputs and speaker vectors for decoding acoustic features.Experimental results have demonstrated that our proposed methods can achieve better objective and subjective performance than the baseline methods without intuitive prosodic features. Besides, the proposed speaker adaption method with utterance-level prosodic features has achieved the best similarity of synthetic speech among all compared methods.

📄 PDF Abstract BibTeX arXiv:2203.00951

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

Controllable speech synthesis by learning discrete phoneme-level prosodic representations

2022-11-29 · Nikolaos Ellinas, Myrsini Christidou, Alexandra Vioni, June Sig Sung 외

In this paper, we present a novel method for phoneme-level prosody control of F0 and duration using intuitive discrete labels. We propose an unsupervised prosodic clustering process which is used to discretize phoneme-le…

ClusteringSpeech Synthesistext-to-speechText to Speech

Prosodic Event Recognition using Convolutional Neural Networks with Context Information

2017-06-02 · Sabrina Stehwien, Ngoc Thang Vu

This paper demonstrates the potential of convolutional neural networks (CNN) for detecting and classifying prosodic events on words, specifically pitch accents and phrase boundary tones, from frame-based acoustic feature…

Position

Controllable neural text-to-speech synthesis using intuitive prosodic features

2020-09-14 · Tuomo Raitio, Ramya Rasipuram, Dan Castellani

Modern neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the prosody of generated utterances often represents the average prosodic style of the database ins…

SentenceSpeech Synthesistext-to-speechText to Speech+1

Improved Prosodic Clustering for Multispeaker and Speaker-independent Phoneme-level Prosody Control

2021-11-19 · Myrsini Christidou, Alexandra Vioni, Nikolaos Ellinas, Georgios Vamvoukakis 외

This paper presents a method for phoneme-level prosody control of F0 and duration on a multispeaker text-to-speech setup, which is based on prosodic clustering. An autoregressive attention-based model is used, incorporat…

ClusteringData Augmentationtext-to-speechText to Speech

Controlling Prosody in End-to-End TTS: A Case Study on Contrastive Focus Generation

2021-11-01 · CoNLL (EMNLP) 2021 11 · Siddique Latif, Inyoung Kim, Ioan Calapodescu, Laurent Besacier

While End-2-End Text-to-Speech (TTS) has made significant progresses over the past few years, these systems still lack intuitive user controls over prosody. For instance, generating speech with fine-grained prosody contr…

text-to-speechText to Speech