paper-with-me

Papers

Unify and Conquer: How Phonetic Feature Representation Affects Polyglot Text-To-Speech (TTS)

2022-07-04 · Ariadna Sanchez, Alessio Falai, Ziyao Zhang, Orazio Angelini, Kayoko Yanagisawa

An essential design decision for multilingual Neural Text-To-Speech (NTTS) systems is how to represent input linguistic features within the model. Looking at the wide variety of approaches in the literature, two main paradigms emerge, unified and separate representations. The former uses a shared set of phonetic tokens across languages, whereas the latter uses unique phonetic tokens for each language. In this paper, we conduct a comprehensive study comparing multilingual NTTS systems models trained with both representations. Our results reveal that the unified approach consistently achieves better cross-lingual synthesis with respect to both naturalness and accent. Separate representations tend to have an order of magnitude more tokens than unified ones, which may affect model capacity. For this reason, we carry out an ablation study to understand the interaction of the representation type with the size of the token embedding. We find that the difference between the two paradigms only emerges above a certain threshold embedding size. This study provides strong evidence that unified representations should be the preferred paradigm when building multilingual NTTS systems.

📄 PDF Abstract BibTeX arXiv:2207.01547

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

Dehumanizing Voice Technology: Phonetic & Experiential Consequences of Restricted Human-Machine Interaction

2021-11-02 · Christian Hildebrand, Donna Hoffman, Tom Novak

The use of natural language and voice-based interfaces gradu-ally transforms how consumers search, shop, and express their preferences. The current work explores how changes in the syntactical structure of the interactio…

The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders

2025-09-19 · Adrian Sauter, Willem Zuidema, Marianne de Heer Kloots arxiv

How does visual information included in training affect language processing in audio- and text-based deep learning models? We explore how such visual grounding affects model-internal representations of words, and find su…

Visual Grounding

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

2026-06-10 · Sangmin Lee, Eekgyun Ahn, Woongjib Choi, Hong-Goo Kang arxiv

We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phoneme (G2P)-based approaches are limited to around 100 languages due to …

Disentangled Phonetic Representation for Chinese Spelling Correction

2023-05-24 · Zihong Liang, Xiaojun Quan, Qifan Wang

Chinese Spelling Correction (CSC) aims to detect and correct erroneous characters in Chinese texts. Although efforts have been made to introduce phonetic information (Hanyu Pinyin) in this task, they typically merge phon…

Spelling Correction

Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis

2025-01-12 · Minu Kim, Kangwook Jang, Hoirin Kim

This paper examines how linguistic similarity affects cross-lingual phonetic representation in speech processing for low-resource languages, emphasizing effective source language selection. Previous cross-lingual researc…

Phoneme RecognitionSelf-Supervised Learning