paper-with-me

홈 › Papers

Fine-grained robust prosody transfer for single-speaker neural text-to-speech

2019-07-04 · Viacheslav Klimkov, Srikanth Ronanki, Jonas Rohnke, Thomas Drugman

We present a neural text-to-speech system for fine-grained prosody transfer from one speaker to another. Conventional approaches for end-to-end prosody transfer typically use either fixed-dimensional or variable-length prosody embedding via a secondary attention to encode the reference signal. However, when trained on a single-speaker dataset, the conventional prosody transfer systems are not robust enough to speaker variability, especially in the case of a reference signal coming from an unseen speaker. Therefore, we propose decoupling of the reference signal alignment from the overall system. For this purpose, we pre-compute phoneme-level time stamps and use them to aggregate prosodic features per phoneme, injecting them into a sequence-to-sequence text-to-speech system. We incorporate a variational auto-encoder to further enhance the latent representation of prosody embeddings. We show that our proposed approach is significantly more stable and achieves reliable prosody transplantation from an unseen speaker. We also propose a solution to the use case in which the transcription of the reference signal is absent. We evaluate all our proposed methods using both objective and subjective listening tests.

📄 PDF Abstract BibTeX arXiv:1907.02479

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer

2022-06-27 · Sri Karlapati, Penny Karanasou, Mateusz Lajszczak, Ammar Abbas 외

In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually appropriate prosody, and c) transferring…

eCat: An End-to-End Model for Multi-Speaker TTS & Many-to-Many Fine-Grained Prosody Transfer

2023-06-20 · Ammar Abbas, Sri Karlapati, Bastian Schnell, Penny Karanasou 외

We present eCat, a novel end-to-end multispeaker model capable of: a) generating long-context speech with expressive and contextually appropriate prosody, and b) performing fine-grained prosody transfer between any pair …

CopyCat: Many-to-Many Fine-Grained Prosody Transfer for Neural Text-to-Speech

2020-04-30

Prosody Transfer (PT) is a technique that aims to use the prosody from a source audio as a reference while synthesising speech. Fine-grained PT aims at capturing prosodic aspects like rhythm, emphasis, melody, duration, …

Rhythmtext-to-speechText to Speech

Cross-speaker Style Transfer with Prosody Bottleneck in Neural Speech Synthesis

2021-07-27 · Shifeng Pan, Lei He

Cross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expressing all styles and to collect correspon…

Expressive Speech SynthesisSpeech SynthesisStyle Transfertext-to-speech+1

Multi-speaker Multi-style Text-to-speech Synthesis With Single-speaker Single-style Training Data Scenarios

2021-12-23 · Qicong Xie, Tao Li, Xinsheng Wang, Zhichao Wang 외

In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker. However, it is hard for one speaker to express all expected styles. …

DiversitySpeech SynthesisStyle Transfertext-to-speech+2