Prosody Prediction
1개 벤치마크 · 논문 16편 · 이 태스크의 논문 보기 →
Benchmarks
Helsinki Prosody Corpus
Most implemented
PRESENT: Zero-Shot Text-to-Prosody Control
Prosody Analysis of Audiobooks
On the Utility of Self-supervised Models for Prosody-related Tasks
Papers
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
Speech synthesis from intracranial EEG (iEEG) signals offers a promising avenue for restoring communication in individuals with severe speech impairments. However, achieving intelligible and natural speech remains challe…
Prosody PredictionSpeech SynthesisVisualSpeech: Enhance Prosody with Visual Context in TTS
Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody from a single text input. While previous research has addressed this by predicting prosodic informatio…
Prosody Predictiontext-to-speechText to SpeechDiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model ba…
Prosody Predictiontext-to-speechText to SpeechWord-wise intonation model for cross-language TTS systems
In this paper we propose a word-wise intonation model for Russian language and show how it can be generalized for other languages. The proposed model is suitable for automatic data markup and its extended application to …
Dynamic Time WarpingProsody Predictiontext-to-speechText to SpeechPRESENT: Zero-Shot Text-to-Prosody Control
Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-t…
Prosody PredictionSpeech Synthesistext-to-speechText to SpeechProsody Analysis of Audiobooks
Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…
AttributeLanguage ModelingLanguage ModellingProsody Prediction+2