paper-with-me

Papers

Learning Multilingual Expressive Speech Representation for Prosody Prediction without Parallel Data

2023-06-29 · Jarod Duret, Titouan Parcollet, Yannick Estève

We propose a method for speech-to-speech emotionpreserving translation that operates at the level of discrete speech units. Our approach relies on the use of multilingual emotion embedding that can capture affective information in a language-independent manner. We show that this embedding can be used to predict the pitch and duration of speech units in a target language, allowing us to resynthesize the source speech signal with the same emotional content. We evaluate our approach to English and French speech signals and show that it outperforms a baseline method that does not use emotional information, including when the emotion embedding is extracted from a different language. Even if this preliminary study does not address directly the machine translation issue, our results demonstrate the effectiveness of our approach for cross-lingual emotion preservation in the context of speech resynthesis.

📄 PDF Abstract BibTeX arXiv:2306.17199

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationProsody PredictionResynthesisTranslation

Similar Papers 제목 키워드 기반

Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages

2026-08-28 · Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman 외 arxiv

Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expressive speech-to-speech translation (S2ST…

Speech-to-Speech Translation

IQDUBBING: Prosody modeling based on discrete self-supervised speech representation for expressive voice conversion

2022-01-02 · Wendong Gan, Bolong Wen, Ying Yan, Haitao Chen 외

Prosody modeling is important, but still challenging in expressive voice conversion. As prosody is difficult to model, and other factors, e.g., speaker, environment and content, which are entangled with prosody in speech…

QuantizationVoice Conversion

Cross-lingual Prosody Transfer for Expressive Machine Dubbing

2023-06-20 · Jakub Swiatkowski, Duo Wang, Mikolaj Babianski, Patrick Lumban Tobing 외

Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to l…

Expressive Speech SynthesisSpeech Synthesis

Ensemble prosody prediction for expressive speech synthesis

2023-04-03 · Tian Huey Teh, Vivian Hu, Devang S Ram Mohan, Zack Hodari 외

Generating expressive speech with rich and varied prosody continues to be a challenge for Text-to-Speech. Most efforts have focused on sophisticated neural architectures intended to better model the data distribution. Ye…

DiversityEnsemble LearningExpressive Speech SynthesisPrediction+4

HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech

2025-09-25 · Aurosweta Mahapatra, Ismail Rasim Ulgen, Berrak Sisman arxiv

Current anti-spoofing systems remain vulnerable to expressive and emotional synthetic speech, since they rarely leverage prosody as a discriminative cue. Prosody is central to human expressiveness and emotion, and humans…

Self-Supervised LearningMulti-Task LearningSpoof Detection