paper-with-me

홈 › Papers

Perception of prosodic variation for speech synthesis using an unsupervised discrete representation of F0

2020-03-14 · Zack Hodari, Catherine Lai, Simon King

In English, prosody adds a broad range of information to segment sequences, from information structure (e.g. contrast) to stylistic variation (e.g. expression of emotion). However, when learning to control prosody in text-to-speech voices, it is not clear what exactly the control is modifying. Existing research on discrete representation learning for prosody has demonstrated high naturalness, but no analysis has been performed on what these representations capture, or if they can generate meaningfully-distinct variants of an utterance. We present a phrase-level variational autoencoder with a multi-modal prior, using the mode centres as "intonation codes". Our evaluation establishes which intonation codes are perceptually distinct, finding that the intonation codes from our multi-modal latent model were significantly more distinct than a baseline using k-means clustering. We carry out a follow-up qualitative study to determine what information the codes are carrying. Most commonly, listeners commented on the intonation codes having a statement or question style. However, many other affect-related styles were also reported, including: emotional, uncertain, surprised, sarcastic, passive aggressive, and upset.

📄 PDF Abstract BibTeX arXiv:2003.06686

Code (1)

ZackHodari/discrete_intonation 공식 구현 pytorch

Tasks

ClusteringRepresentation LearningSpeech Synthesistext-to-speechText to Speech

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Prosodic Clustering for Phoneme-level Prosody Control in End-to-End Speech Synthesis

2021-11-19 · Alexandra Vioni, Myrsini Christidou, Nikolaos Ellinas, Georgios Vamvoukakis 외

This paper presents a method for controlling the prosody at the phoneme level in an autoregressive attention-based text-to-speech system. Instead of learning latent prosodic features with a variational framework as is co…

ClusteringDecoderSpeech Synthesistext-to-speech+1

Prosodic Prominence and Boundaries in Sequence-to-Sequence Speech Synthesis

2020-06-29

Recent advances in deep learning methods have elevated synthetic speech quality to human level, and the field is now moving towards addressing prosodic variation in synthetic speech.Despite successes in this effort, the …

SentenceSpeech Synthesis

Uncovering Latent Style Factors for Expressive Speech Synthesis

2017-11-01 · Yuxuan Wang, RJ Skerry-Ryan, Ying Xiao, Daisy Stanton 외

Prosodic modeling is a core problem in speech synthesis. The key challenge is producing desirable prosody from textual input containing only phonetic information. In this preliminary study, we introduce the concept of "s…

Expressive Speech SynthesisSpeech Synthesis

Controllable neural text-to-speech synthesis using intuitive prosodic features

2020-09-14 · Tuomo Raitio, Ramya Rasipuram, Dan Castellani

Modern neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the prosody of generated utterances often represents the average prosodic style of the database ins…

SentenceSpeech Synthesistext-to-speechText to Speech+1

Variations prosodiques en synth\`ese par s\'election d'unit\'es: l'exemple des phrases interrogatives (Prosodic variations in unit-based speech synthesis: the example of interrogative sentences) [in French]

2012-06-01 · JEPTALNRECITAL 2012 6 · Laurence Martin, Sophie Roekhaut, Richard Beaufort
Speech SynthesisText-To-Speech Synthesis