paper-with-me

Papers

VisualSpeech: Enhance Prosody with Visual Context in TTS

2025-01-31 · Shumin Que, Anton Ragni

Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody from a single text input. While previous research has addressed this by predicting prosodic information from both text and speech, additional contextual information, such as visual features, remains underutilized. This paper investigates the potential of integrating visual context to enhance prosody prediction. We propose a novel model, VisualSpeech, which incorporates both visual and textual information for improved prosody generation. Empirical results demonstrate that visual features provide valuable prosodic cues beyond the textual input, significantly enhancing the naturalness and accuracy of the synthesized speech. Audio samples are available at https://ariameetgit.github.io/VISUALSPEECH-SAMPLES/.

📄 PDF Abstract BibTeX arXiv:2501.19258

Code (0)

등록된 구현이 없습니다.

Tasks

Prosody Predictiontext-to-speechText to Speech

Similar Papers 제목 키워드 기반

MCDubber: Multimodal Context-Aware Expressive Video Dubbing

2024-08-21 · Yuan Zhao, Zhenqi Jia, Rui Liu, De Hu 외

Automatic Video Dubbing (AVD) aims to take the given script and generate speech that aligns with lip motion and prosody expressiveness. Current AVD models mainly utilize visual information of the current sentence to enha…

Sentence

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

2024-12-25 · Yuan Zhao, Rui Liu, Gaoxiang Cong

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issue…

Graph AttentionSentence

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

2025-02-27 · Weihao wu, Zhiwei Lin, Yixuan Zhou, Jingbei Li 외

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, exist…

DiversityLanguage ModelingLanguage ModellingSpeech Synthesis

Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing

2025-03-15 · CVPR 2025 1 · Zhedong Zhang, Liang Li, Chenggang Yan, Chunshan Liu 외

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip.…

Emotion Recognition

Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

2025-09-07 · Zhenqi Jia, Rui Liu, Berrak Sisman, Haizhou Li arxiv

Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate prosody expression of the target utterance …

Speech Synthesis