paper-with-me

홈 › Papers

eCat: An End-to-End Model for Multi-Speaker TTS & Many-to-Many Fine-Grained Prosody Transfer

2023-06-20 · Ammar Abbas, Sri Karlapati, Bastian Schnell, Penny Karanasou, Marcel Granero Moya, Amith Nagaraj, Ayman Boustati, Nicole Peinelt, Alexis Moinet, Thomas Drugman

We present eCat, a novel end-to-end multispeaker model capable of: a) generating long-context speech with expressive and contextually appropriate prosody, and b) performing fine-grained prosody transfer between any pair of seen speakers. eCat is trained using a two-stage training approach. In Stage I, the model learns speaker-independent word-level prosody representations in an end-to-end fashion from speech. In Stage II, we learn to predict the prosody representations using the contextual information available in text. We compare eCat to CopyCat2, a model capable of both fine-grained prosody transfer (FPT) and multi-speaker TTS. We show that eCat statistically significantly reduces the gap in naturalness between CopyCat2 and human recordings by an average of 46.7% across 2 languages, 3 locales, and 7 speakers, along with better target-speaker similarity in FPT. We also compare eCat to VITS, and show a statistically significant preference.

📄 PDF Abstract BibTeX arXiv:2306.11327

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An LSTM-Based Deep Learning Approach for Detecting Self-Deprecating Sarcasm in Textual Data

2019-12-01 · ICON 2019 12 · Ashraf Kamal, Muhammad Abulaish

Self-deprecating sarcasm is a special category of sarcasm, which is nowadays popular and useful for many real-life applications, such as brand endorsement, product campaign, digital marketing, and advertisement. The self…

Deep LearningMarketingSarcasm Detection

CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer

2022-06-27 · Sri Karlapati, Penny Karanasou, Mateusz Lajszczak, Ammar Abbas 외

In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually appropriate prosody, and c) transferring…

ECAT: A Entire space Continual and Adaptive Transfer Learning Framework for Cross-Domain Recommendation

2024-07-02 · Chaoqun Hou, Yuanhang Zhou, Yi Cao, Tong Liu

In industrial recommendation systems, there are several mini-apps designed to meet the diverse interests and needs of users. The sample space of them is merely a small subset of the entire space, making it challenging to…

Domain AdaptationKnowledge DistillationRecommendation SystemsTransfer Learning

CopyCat: Many-to-Many Fine-Grained Prosody Transfer for Neural Text-to-Speech

2020-04-30

Prosody Transfer (PT) is a technique that aims to use the prosody from a source audio as a reference while synthesising speech. Fine-grained PT aims at capturing prosodic aspects like rhythm, emphasis, melody, duration, …

Rhythmtext-to-speechText to Speech

Many-to-Many Voice Conversion based Feature Disentanglement using Variational Autoencoder

2021-07-11 · Manh Luong, Viet Anh Tran

Voice conversion is a challenging task which transforms the voice characteristics of a source speaker to a target speaker without changing linguistic content. Recently, there have been many works on many-to-many Voice Co…

DisentanglementVoice Conversion