paper-with-me

Papers Prosody Prediction

“Prosody Prediction” 태그가 달린 논문 16편 · 필터 해제

MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction

2025-08-05 · Mohammed Salah Al-Radhi, Géza Németh, Branislav Gerazov arxiv

Speech synthesis from intracranial EEG (iEEG) signals offers a promising avenue for restoring communication in individuals with severe speech impairments. However, achieving intelligible and natural speech remains challe…

Prosody PredictionSpeech Synthesis

VisualSpeech: Enhance Prosody with Visual Context in TTS

2025-01-31 · Shumin Que, Anton Ragni

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody from a single text input. While previous research has addressed this by predicting prosodic informatio…

Prosody Predictiontext-to-speechText to Speech

DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles

2024-12-04 · Jiaxuan Liu, Zhaoci Liu, Yajun Hu, Yingying Gao 외

Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model ba…

Prosody Predictiontext-to-speechText to Speech

Word-wise intonation model for cross-language TTS systems

2024-09-30 · Tomilov A. A., Gromova A. Y., Svischev A. N

In this paper we propose a word-wise intonation model for Russian language and show how it can be generalized for other languages. The proposed model is suitable for automatic data markup and its extended application to …

Dynamic Time WarpingProsody Predictiontext-to-speechText to Speech

PRESENT: Zero-Shot Text-to-Prosody Control

2024-08-13 · Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman 외

Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-t…

Prosody PredictionSpeech Synthesistext-to-speechText to Speech

Prosody Analysis of Audiobooks

2023-10-10 · Charuta Pethe, Bach Pham, Felix D Childress, Yunting Yin 외

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…

AttributeLanguage ModelingLanguage ModellingProsody Prediction+2

A Comparative Analysis of Pretrained Language Models for Text-to-Speech

2023-09-04 · Marcel Granero-Moya, Penny Karanasou, Sri Karlapati, Bastian Schnell 외

State-of-the-art text-to-speech (TTS) systems have utilized pretrained language models (PLMs) to enhance prosody and create more natural-sounding speech. However, while PLMs have been extensively researched for natural l…

Natural Language UnderstandingPredictionProsody Predictiontext-to-speech+1

Learning Multilingual Expressive Speech Representation for Prosody Prediction without Parallel Data

2023-06-29 · Jarod Duret, Titouan Parcollet, Yannick Estève

We propose a method for speech-to-speech emotionpreserving translation that operates at the level of discrete speech units. Our approach relies on the use of multilingual emotion embedding that can capture affective info…

Machine TranslationProsody PredictionResynthesisTranslation

What Can an Accent Identifier Learn? Probing Phonetic and Prosodic Information in a Wav2vec2-based Accent Identification Model

2023-06-10 · Mu Yang, Ram C. M. C. Shekar, Okim Kang, John H. L. Hansen

This study is focused on understanding and quantifying the change in phoneme and prosody information encoded in the Self-Supervised Learning (SSL) model, brought by an accent identification (AID) fine-tuning task. This p…

Automatic Speech RecognitionProsody PredictionSelf-Supervised Learningspeech-recognition+1

Ensemble prosody prediction for expressive speech synthesis

2023-04-03 · Tian Huey Teh, Vivian Hu, Devang S Ram Mohan, Zack Hodari 외

Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Ye…

DiversityEnsemble LearningExpressive Speech SynthesisPrediction+4

Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis

2023-03-14 · Chunyu Qiang, Peng Yang, Hao Che, Ying Zhang 외

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker's timbre. In most previous methods, the synthesized fine-grained prosody feature…

Prosody PredictionSpeech SynthesisStyle Transfer

On the Utility of Self-supervised Models for Prosody-related Tasks

2022-10-13 · Guan-Ting Lin, Chi-Luen Feng, Wei-Ping Huang, Yuan Tseng 외

Self-Supervised Learning (SSL) from speech data has produced models that have achieved remarkable performance in many tasks, and that are known to implicitly represent many aspects of information latently present in spee…

Prosody PredictionSelf-Supervised Learning

Prosody Learning Mechanism for Speech Synthesis System Without Text Length Limit

2020-08-13 · Zhen Zeng, Jianzong Wang, Ning Cheng, Jing Xiao

Recent neural speech synthesis systems have gradually focused on the control of prosody to improve the quality of synthesized speech, but they rarely consider the variability of prosody and the correlation between prosod…

Language ModelingLanguage ModellingPositionProsody Prediction+1

Controllable Sequence-To-Sequence Neural TTS with LPCNET Backend for Real-time Speech Synthesis on CPU

2020-02-25

State-of-the-art sequence-to-sequence acoustic networks, that convert a phonetic sequence to a sequence of spectral features with no explicit prosody prediction, generate speech with close to natural quality, when cascad…

CPUProsody PredictionSentenceSpeech Synthesis

Predicting Prosodic Prominence from Text with Pre-trained Contextualized Word Representations

2019-08-06 · WS (NoDaLiDa) 2019 9 · Aarne Talman, Antti Suni, Hande Celikkanat, Sofoklis Kakouros 외

In this paper we introduce a new natural language processing dataset and benchmark for predicting prosodic prominence from written text. To our knowledge this will be the largest publicly available dataset with prosodic …

Prosody Prediction

Automatic Prosody Prediction for Chinese Speech Synthesis using BLSTM-RNN and Embedding Features

2015-11-02 · Chuang Ding, Lei Xie, Jie Yan, Weini Zhang 외

Prosody affects the naturalness and intelligibility of speech. However, automatic prosody prediction from text for Chinese speech synthesis is still a great challenge and the traditional conditional random fields (CRF) b…

Feature EngineeringProsody PredictionSpeech Synthesis
1–16 / 16