paper-with-me

Papers

Do Prosody Transfer Models Transfer Prosody?

2023-03-07 · Atli Thor Sigurgeirsson, Simon King

Some recent models for Text-to-Speech synthesis aim to transfer the prosody of a reference utterance to the generated target synthetic speech. This is done by using a learned embedding of the reference utterance, which is used to condition speech generation. During training, the reference utterance is identical to the target utterance. Yet, during synthesis, these models are often used to transfer prosody from a reference that differs from the text or speaker being synthesized. To address this inconsistency, we propose to use a different, but prosodically-related, utterance during training too. We believe this should encourage the model to learn to transfer only those characteristics that the reference and target have in common. If prosody transfer methods do indeed transfer prosody they should be able to be trained in the way we propose. However, results show that a model trained under these conditions performs significantly worse than one trained using the target utterance as a reference. To explain this, we hypothesize that prosody transfer models do not learn a transferable representation of prosody, but rather an utterance-level representation which is highly dependent on both the reference speaker and reference text.

📄 PDF Abstract BibTeX arXiv:2303.04289

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer

2022-06-27 · Sri Karlapati, Penny Karanasou, Mateusz Lajszczak, Ammar Abbas 외

In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually appropriate prosody, and c) transferring…

Global Rhythm Style Transfer Without Text Transcriptions

2021-06-16 · Kaizhi Qian, Yang Zhang, Shiyu Chang, JinJun Xiong 외

Prosody plays an important role in characterizing the style of a speaker or an emotion, but most non-parallel voice or emotion style transfer algorithms do not convert any prosody information. Two major components of pro…

Representation LearningRhythmStyle Transfer

Cross-lingual Prosody Transfer for Expressive Machine Dubbing

2023-06-20 · Jakub Swiatkowski, Duo Wang, Mikolaj Babianski, Patrick Lumban Tobing 외

Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to l…

Expressive Speech SynthesisSpeech Synthesis

ADEPT: A Dataset for Evaluating Prosody Transfer

2021-06-15 · Alexandra Torresquintero, Tian Huey Teh, Christopher G. R. Wallis, Marlene Staib 외

Text-to-speech is now able to achieve near-human naturalness and research focus has shifted to increasing expressivity. One popular method is to transfer the prosody from a reference speech sample. There have been consid…

text-to-speechText to Speech

Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer

2023-06-20 · Jakub Swiatkowski, Duo Wang, Mikolaj Babianski, Giuseppe Coccia 외

Speech generation for machine dubbing adds complexity to conventional Text-To-Speech solutions as the generated output is required to match the expressiveness, emotion and speaking rate of the source content. Capturing a…

text-to-speechText to Speech