paper-with-me

Prosody Prediction

1개 벤치마크 · 논문 16편 · 이 태스크의 논문 보기 →

Benchmarks

Most implemented

PRESENT: Zero-Shot Text-to-Prosody Control

2024-08-13 · 구현 1개

Prosody Analysis of Audiobooks

2023-10-10 · 구현 1개

Papers

MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction

2025-08-05 · Mohammed Salah Al-Radhi, Géza Németh, Branislav Gerazov arxiv

Speech synthesis from intracranial EEG (iEEG) signals offers a promising avenue for restoring communication in individuals with severe speech impairments. However, achieving intelligible and natural speech remains challe…

Prosody PredictionSpeech Synthesis

VisualSpeech: Enhance Prosody with Visual Context in TTS

2025-01-31 · Shumin Que, Anton Ragni

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody from a single text input. While previous research has addressed this by predicting prosodic informatio…

Prosody Predictiontext-to-speechText to Speech

DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles

2024-12-04 · Jiaxuan Liu, Zhaoci Liu, Yajun Hu, Yingying Gao 외

Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model ba…

Prosody Predictiontext-to-speechText to Speech

Word-wise intonation model for cross-language TTS systems

2024-09-30 · Tomilov A. A., Gromova A. Y., Svischev A. N

In this paper we propose a word-wise intonation model for Russian language and show how it can be generalized for other languages. The proposed model is suitable for automatic data markup and its extended application to …

Dynamic Time WarpingProsody Predictiontext-to-speechText to Speech

PRESENT: Zero-Shot Text-to-Prosody Control

2024-08-13 · Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman 외

Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-t…

Prosody PredictionSpeech Synthesistext-to-speechText to Speech

Prosody Analysis of Audiobooks

2023-10-10 · Charuta Pethe, Bach Pham, Felix D Childress, Yunting Yin 외

Recent advances in text-to-speech have made it possible to generate natural-sounding audio from text. However, audiobook narrations involve dramatic vocalizations and intonations by the reader, with greater reliance on e…

AttributeLanguage ModelingLanguage ModellingProsody Prediction+2

전체 16편 보기 →