paper-with-me

Papers

METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer

2023-07-29 · Xinfa Zhu, Yi Lei, Tao Li, Yongmao Zhang, Hongbin Zhou, Heng Lu, Lei Xie

Previous multilingual text-to-speech (TTS) approaches have considered leveraging monolingual speaker data to enable cross-lingual speech synthesis. However, such data-efficient approaches have ignored synthesizing emotional aspects of speech due to the challenges of cross-speaker cross-lingual emotion transfer - the heavy entanglement of speaker timbre, emotion, and language factors in the speech signal will make a system produce cross-lingual synthetic speech with an undesired foreign accent and weak emotion expressiveness. This paper proposes the Multilingual Emotional TTS (METTS) model to mitigate these problems, realizing both cross-speaker and cross-lingual emotion transfer. Specifically, METTS takes DelightfulTTS as the backbone model and proposes the following designs. First, to alleviate the foreign accent problem, METTS introduces multi-scale emotion modeling to disentangle speech prosody into coarse-grained and fine-grained scales, producing language-agnostic and language-specific emotion representations, respectively. Second, as a pre-processing step, formant shift-based information perturbation is applied to the reference signal for better disentanglement of speaker timbre in the speech. Third, a vector quantization-based emotion matcher is designed for reference selection, leading to decent naturalness and emotion diversity in cross-lingual synthetic speech. Experiments demonstrate the good design of METTS.

📄 PDF Abstract BibTeX arXiv:2307.15951

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementDiversityQuantizationSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

2024-04-29 · Zhi-Qi Cheng, Xiang Li, Jun-Yan He, Junyao Chen 외

Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture …

Contrastive LearningSpeech Synthesistext-to-speechText to Speech+1

Optimizing Multilingual Text-To-Speech with Accents & Emotions

2025-06-19 · Pranav Pawar, Akshansh Dwivedi, Jenish Boricha, Himanshu Gohil 외

State-of-the-art text-to-speech (TTS) systems realize high naturalness in monolingual environments, synthesizing speech with correct multilingual accents (especially for Indic languages) and context-relevant emotions sti…

DisentanglementEmotion Recognitiontext-to-speechText to Speech+1

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

2025-03-25 · Heqing Zou, Fengmao Lv, Desheng Zheng, Eng Siong Chng 외

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses sig…

Contrastive LearningDiversityEmotion RecognitionSpeech Emotion Recognition

CAMEO: Collection of Multilingual Emotional Speech Corpora

2025-05-16 · Iwona Christop, Maciej Czajka

This paper presents CAMEO -- a curated collection of multilingual emotional speech datasets designed to facilitate research in emotion recognition and other speech-related tasks. The main objectives were to ensure easy a…

Emotion RecognitionSpeech Emotion Recognition

Learning Multilingual Expressive Speech Representation for Prosody Prediction without Parallel Data

2023-06-29 · Jarod Duret, Titouan Parcollet, Yannick Estève

We propose a method for speech-to-speech emotionpreserving translation that operates at the level of discrete speech units. Our approach relies on the use of multilingual emotion embedding that can capture affective info…

Machine TranslationProsody PredictionResynthesisTranslation