Learning Multilingual Expressive Speech Representation for Prosody Prediction without Parallel Data
We propose a method for speech-to-speech emotionpreserving translation that operates at the level of discrete speech units. Our approach relies on the use of multilingual emotion embedding that can capture affective information in a language-independent manner. We show that this embedding can be used to predict the pitch and duration of speech units in a target language, allowing us to resynthesize the source speech signal with the same emotional content. We evaluate our approach to English and French speech signals and show that it outperforms a baseline method that does not use emotional information, including when the emotion embedding is extracted from a different language. Even if this preliminary study does not address directly the machine translation issue, our results demonstrate the effectiveness of our approach for cross-lingual emotion preservation in the context of speech resynthesis.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationProsody PredictionResynthesisTranslationSimilar Papers 제목 키워드 기반
Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages
Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expressive speech-to-speech translation (S2ST…
Speech-to-Speech TranslationIQDUBBING: Prosody modeling based on discrete self-supervised speech representation for expressive voice conversion
Prosody modeling is important, but still challenging in expressive voice conversion. As prosody is difficult to model, and other factors, e.g., speaker, environment and content, which are entangled with prosody in speech…
QuantizationVoice ConversionCross-lingual Prosody Transfer for Expressive Machine Dubbing
Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to l…
Expressive Speech SynthesisSpeech SynthesisEnsemble prosody prediction for expressive speech synthesis
Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Ye…
DiversityEnsemble LearningExpressive Speech SynthesisPrediction+4HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
Current anti-spoofing systems remain vulnerable to expressive and emotional synthetic speech, since they rarely leverage prosody as a discriminative cue. Prosody is central to human expressiveness and emotion, and humans…
Self-Supervised LearningMulti-Task LearningSpoof Detection