QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis
Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and control intonation to further deliver the speaker's questioning intention while transferring emotion from reference speech. We propose a multi-style extractor to extract style embedding from two different levels. While the sentence level represents emotion, the final syllable level represents intonation. For fine-grained intonation control, we use relative attributes to represent intonation intensity at the syllable level.Experiments have validated the effectiveness of QI-TTS for improving intonation expressiveness in emotional speech synthesis.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotional Speech SynthesisSentenceSpeech Synthesistext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Perception of prosodic variation for speech synthesis using an unsupervised discrete representation of F0
In English, prosody adds a broad range of information to segment sequences, from information structure (e.g. contrast) to stylistic variation (e.g. expression of emotion). However, when learning to control prosody in tex…
ClusteringRepresentation LearningSpeech Synthesistext-to-speech+1EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human …
text-to-speechText to SpeechPenambahan emosi menggunakan metode manipulasi prosodi untuk sistem text to speech bahasa Indonesia
Adding an emotions using prosody manipulation method for Indonesian text to speech system. Text To Speech (TTS) is a system that can convert text in one language into speech, accordance with the reading of the text in th…
Sentencetext-to-speechText to SpeechProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
Prosody contains rich information beyond the literal meaning of words, which is crucial for the intelligibility of speech. Current models still fall short in phrasing and intonation; they not only miss or misplace breaks…
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisEmoSpeech: Guiding FastSpeech2 Towards Emotional Text to Speech
State-of-the-art speech synthesis models try to get as close as possible to the human voice. Hence, modelling emotions is an essential part of Text-To-Speech (TTS) research. In our work, we selected FastSpeech2 as the st…
Emotion RecognitionSpeech Synthesistext-to-speechText to Speech