Papers Prosody Prediction
“Prosody Prediction” 태그가 달린 논문 16편 · 필터 해제
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
Speech synthesis from intracranial EEG (iEEG) signals offers a promising avenue for restoring communication in individuals with severe speech impairments. However, achieving intelligible and natural speech remains challe…
Prosody PredictionSpeech SynthesisVisualSpeech: Enhance Prosody with Visual Context in TTS
Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody from a single text input. While previous research has addressed this by predicting prosodic informatio…
Prosody Predictiontext-to-speechText to SpeechDiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model ba…
Prosody Predictiontext-to-speechText to SpeechWord-wise intonation model for cross-language TTS systems
In this paper we propose a word-wise intonation model for Russian language and show how it can be generalized for other languages. The proposed model is suitable for automatic data markup and its extended application to …
Dynamic Time WarpingProsody Predictiontext-to-speechText to SpeechPRESENT: Zero-Shot Text-to-Prosody Control
Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-t…
Prosody PredictionSpeech Synthesistext-to-speechText to SpeechProsody Analysis of Audiobooks
Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…
AttributeLanguage ModelingLanguage ModellingProsody Prediction+2A Comparative Analysis of Pretrained Language Models for Text-to-Speech
State-of-the-art text-to-speech (TTS) systems have utilized pretrained language models (PLMs) to enhance prosody and create more natural-sounding speech. However, while PLMs have been extensively researched for natural l…
Natural Language UnderstandingPredictionProsody Predictiontext-to-speech+1Learning Multilingual Expressive Speech Representation for Prosody Prediction without Parallel Data
We propose a method for speech-to-speech emotionpreserving translation that operates at the level of discrete speech units. Our approach relies on the use of multilingual emotion embedding that can capture affective info…
Machine TranslationProsody PredictionResynthesisTranslationWhat Can an Accent Identifier Learn? Probing Phonetic and Prosodic Information in a Wav2vec2-based Accent Identification Model
This study is focused on understanding and quantifying the change in phoneme and prosody information encoded in the Self-Supervised Learning (SSL) model, brought by an accent identification (AID) fine-tuning task. This p…
Automatic Speech RecognitionProsody PredictionSelf-Supervised Learningspeech-recognition+1Ensemble prosody prediction for expressive speech synthesis
Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Ye…
DiversityEnsemble LearningExpressive Speech SynthesisPrediction+4Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis
Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker's timbre. In most previous methods, the synthesized fine-grained prosody feature…
Prosody PredictionSpeech SynthesisStyle TransferOn the Utility of Self-supervised Models for Prosody-related Tasks
Self-Supervised Learning (SSL) from speech data has produced models that have achieved remarkable performance in many tasks, and that are known to implicitly represent many aspects of information latently present in spee…
Prosody PredictionSelf-Supervised LearningProsody Learning Mechanism for Speech Synthesis System Without Text Length Limit
Recent neural speech synthesis systems have gradually focused on the control of prosody to improve the quality of synthesized speech, but they rarely consider the variability of prosody and the correlation between prosod…
Language ModelingLanguage ModellingPositionProsody Prediction+1Controllable Sequence-To-Sequence Neural TTS with LPCNET Backend for Real-time Speech Synthesis on CPU
State-of-the-art sequence-to-sequence acoustic networks, that convert a phonetic sequence to a sequence of spectral features with no explicit prosody prediction, generate speech with close to natural quality, when cascad…
CPUProsody PredictionSentenceSpeech SynthesisPredicting Prosodic Prominence from Text with Pre-trained Contextualized Word Representations
In this paper we introduce a new natural language processing dataset and benchmark for predicting prosodic prominence from written text. To our knowledge this will be the largest publicly available dataset with prosodic …
Prosody PredictionAutomatic Prosody Prediction for Chinese Speech Synthesis using BLSTM-RNN and Embedding Features
Prosody affects the naturalness and intelligibility of speech. However, automatic prosody prediction from text for Chinese speech synthesis is still a great challenge and the traditional conditional random fields (CRF) b…
Feature EngineeringProsody PredictionSpeech Synthesis