paper-with-me

홈 › Papers

Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control

2024-09-26 · Ryuichi Yamamoto, Yuma Shirahata, Masaya Kawamura, Kentaro Tachibana

We propose a novel description-based controllable text-to-speech (TTS) method with cross-lingual control capability. To address the lack of audio-description paired data in the target language, we combine a TTS model trained on the target language with a description control model trained on another language, which maps input text descriptions to the conditional features of the TTS model. These two models share disentangled timbre and style representations based on self-supervised learning (SSL), allowing for disentangled voice control, such as controlling speaking styles while retaining the original timbre. Furthermore, because the SSL-based timbre and style representations are language-agnostic, combining the TTS and description control models while sharing the same embedding space effectively enables cross-lingual control of voice characteristics. Experiments on English and Japanese TTS demonstrate that our method achieves high naturalness and controllability for both languages, even though no Japanese audio-description pairs are used.

📄 PDF Abstract BibTeX arXiv:2409.17452

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learningtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations

2025-05-24 · Ashwin Sankar, Yoach Lacombe, Sherry Thomas, Praveen Srinivasa Varadhan 외

We introduce RASMALAI, a large-scale speech dataset with rich text descriptions, designed to advance controllable and expressive text-to-speech (TTS) synthesis for 23 Indian languages and English. It comprises 13,000 hou…

Expressive Speech SynthesisSpeech Synthesistext-to-speechText to Speech

VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech

2024-06-12 · Ashishkumar Gudmalwar, Nirmesh Shah, Sai Akarsh, Pankaj Wasnik 외

Despite the significant advancements in Text-to-Speech (TTS) systems, their full utilization in automatic dubbing remains limited. This task necessitates the extraction of voice identity and emotional style from a refere…

text-to-speechText to Speech

Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic Posteriorgram

2021-02-03 · Shengkui Zhao, Hao Wang, Trung Hieu Nguyen, Bin Ma

Cross-lingual voice conversion (VC) is an important and challenging problem due to significant mismatches of the phonetic set and the speech prosody of different languages. In this paper, we build upon the neural text-to…

text-to-speechText to SpeechVoice Conversion

Qwen3-TTS Technical Report

2026-01-22 · Hangrui Hu, Xinfa Zhu, Ting He, Dake Guo 외 arxiv

In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-ba…

PromptTTS: Controllable Text-to-Speech with Text Descriptions

2022-11-22 · Zhifang Guo, Yichong Leng, Yihan Wu, Sheng Zhao 외

Using a text description as prompt to guide the generation of text or images (e.g., GPT-3 or DALLE-2) has drawn wide attention recently. Beyond text and image generation, in this work, we explore the possibility of utili…

DecoderSpeech Synthesistext-to-speechText to Speech