paper-with-me

홈 › Papers

Mix and Match: An Empirical Study on Training Corpus Composition for Polyglot Text-To-Speech (TTS)

2022-07-04 · Ziyao Zhang, Alessio Falai, Ariadna Sanchez, Orazio Angelini, Kayoko Yanagisawa

Training multilingual Neural Text-To-Speech (NTTS) models using only monolingual corpora has emerged as a popular way for building voice cloning based Polyglot NTTS systems. In order to train these models, it is essential to understand how the composition of the training corpora affects the quality of multilingual speech synthesis. In this context, it is common to hear questions such as "Would including more Spanish data help my Italian synthesis, given the closeness of both languages?". Unfortunately, we found existing literature on the topic lacking in completeness in this regard. In the present work, we conduct an extensive ablation study aimed at understanding how various factors of the training corpora, such as language family affiliation, gender composition, and the number of speakers, contribute to the quality of Polyglot synthesis. Our findings include the observation that female speaker data are preferred in most scenarios, and that it is not always beneficial to have more speakers from the target language variant in the training corpus. The findings herein are informative for the process of data procurement and corpora building.

📄 PDF Abstract BibTeX arXiv:2207.01507

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to SpeechVoice Cloning

Similar Papers 제목 키워드 기반

Is Japanese CCGBank empirically correct? A case study of passive and causative constructions

2023-02-28 · Daisuke Bekki, Hitomi Yanaka

The Japanese CCGBank serves as training and evaluation data for developing Japanese CCG parsers. However, since it is automatically generated from the Kyoto Corpus, a dependency treebank, its linguistic validity still ne…

Semantic Parsing

On the Emergence and Test-Time Use of Structural Information in Large Language Models

2026-01-25 · Michelle Chao Chen, Moritz Miller, Bernhard Schölkopf, Siyuan Guo arxiv

Learning structural information from observational data is central to producing new knowledge outside the training corpus. This holds for mechanistic understanding in scientific discovery as well as flexible test-time co…

Mix, Don't Pick: Why Synthetic Corpus Composition Matters for Time Series Foundation Model Pretraining

2026-06-06 · Aaryan Nagpal, Debdeep Sanyal, Murari Mandal, Dhruv Kumar 외 arxiv

Choosing the wrong synthetic generator for time-series foundation model pretraining is costly: under identical training budgets, the best and worst generators produce up to a $2\times$ gap in forecasting error, yet the f…

Does Continued Pretraining on a Learner Corpus Improve Automated Essay Scoring on English Proficiency Tests? Evidence from EFCAMDAT

2026-05-25 · Duy Anh Nguyen arxiv

Automated Essay Scoring (AES) for English proficiency assessment increasingly relies on pretrained transformer models, yet these models are typically trained on general-domain English and may under-represent second-langu…

Automated Essay Scoring

A Study of Lagrangean Decompositions and Dual Ascent Solvers for Graph Matching

2016-12-16 · CVPR 2017 7 · Paul Swoboda, Carsten Rother, Hassan Abu Alhaija, Dagmar Kainmueller 외

We study the quadratic assignment problem, in computer vision also known as graph matching. Two leading solvers for this problem optimize the Lagrange decomposition duals with sub-gradient and dual ascent (also known as …

Graph Matching