paper-with-me

Papers

XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

2024-06-07 · Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar, Iulian Gulea, Logan Hart, Aya Aljafari, Joshua Meyer, Reuben Morais, Samuel Olayemi, Julian Weber

Most Zero-shot Multi-speaker TTS (ZS-TTS) systems support only a single language. Although models like YourTTS, VALL-E X, Mega-TTS 2, and Voicebox explored Multilingual ZS-TTS they are limited to just a few high/medium resource languages, limiting the applications of these models in most of the low/medium resource languages. In this paper, we aim to alleviate this issue by proposing and making publicly available the XTTS system. Our method builds upon the Tortoise model and adds several novel modifications to enable multilingual training, improve voice cloning, and enable faster training and inference. XTTS was trained in 16 languages and achieved state-of-the-art (SOTA) results in most of them.

📄 PDF Abstract BibTeX arXiv:2406.04904

Code (1)

Edresson/ZS-TTS-Evaluation 공식 구현 pytorch

Tasks

text-to-speechText to SpeechVoice CloningZero-Shot Multi-Speaker TTS

Similar Papers 제목 키워드 기반

IndexTTS 2.5 Technical Report

2026-01-07 · Yunpei Li, Xun Zhou, Jinchao Wang, Lu Wang 외 arxiv

In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic (T2S) module and a non-autoregressive Semantic-to-Mel (S2M)…

Reinforcement Learning

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

2025-02-08 · Wei Deng, Siyi Zhou, Jingchen Shu, Jinchao Wang 외

Recently, large language model (LLM) based text-to-speech (TTS) systems have gradually become the mainstream in the industry due to their high naturalness and powerful zero-shot voice cloning capabilities.Here, we introd…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+4

Towards Zero-Shot Text-To-Speech for Arabic Dialects

2024-06-24 · Khai Duy Doan, Abdul Waheed, Muhammad Abdul-Mageed

Zero-shot multi-speaker text-to-speech (ZS-TTS) systems have advanced for English, however, it still lags behind due to insufficient resources. We address this gap for Arabic, a language of more than 450 million native s…

Dialect IdentificationSpeech Synthesistext-to-speechText to Speech

Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation

2020-04-24 · ACL 2020 6 · Biao Zhang, Philip Williams, Ivan Titov, Rico Sennrich

Massively multilingual models for neural machine translation (NMT) are theoretically attractive, but often underperform bilingual models and deliver poor zero-shot translations. In this paper, we explore ways to improve …

Machine TranslationNMTTranslation

Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations

2023-06-14 · Gregor Geigle, Radu Timofte, Goran Glavaš

Vision-and-language (VL) models with separate encoders for each modality (e.g., CLIP) have become the go-to models for zero-shot image classification and image-text retrieval. They are, however, mostly evaluated in Engli…

image-classificationImage ClassificationImage-text RetrievalMachine Translation+3