VisualSpeech: Enhance Prosody with Visual Context in TTS
Text-to-Speech (TTS) synthesis faces the inherent challenge of producing multiple speech outputs with varying prosody from a single text input. While previous research has addressed this by predicting prosodic information from both text and speech, additional contextual information, such as visual features, remains underutilized. This paper investigates the potential of integrating visual context to enhance prosody prediction. We propose a novel model, VisualSpeech, which incorporates both visual and textual information for improved prosody generation. Empirical results demonstrate that visual features provide valuable prosodic cues beyond the textual input, significantly enhancing the naturalness and accuracy of the synthesized speech. Audio samples are available at https://ariameetgit.github.io/VISUALSPEECH-SAMPLES/.
Code (0)
등록된 구현이 없습니다.
Tasks
Prosody Predictiontext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
MCDubber: Multimodal Context-Aware Expressive Video Dubbing
Automatic Video Dubbing (AVD) aims to take the given script and generate speech that aligns with lip motion and prosody expressiveness. Current AVD models mainly utilize visual information of the current sentence to enha…
SentenceTowards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issue…
Graph AttentionSentenceDiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, exist…
DiversityLanguage ModelingLanguage ModellingSpeech SynthesisProsody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip.…
Emotion RecognitionMultimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis
Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate prosody expression of the target utterance …
Speech Synthesis