paper-with-me

Papers

Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis

2026-01-20 · Thanathai Lertpetchpun, Yoonjeong Lee, Thanapat Trachu, Jihwan Lee, Tiantian Feng, Dani Byrd, Shrikanth Narayanan arxiv

Many spoken languages, including English, exhibit wide variation in dialects and accents, making accent control an important capability for flexible text-to-speech (TTS) models. Current TTS systems typically generate accented speech by conditioning on speaker embeddings associated with specific accents. While effective, this approach offers limited interpretability and controllability, as embeddings also encode traits such as timbre and emotion. In this study, we analyze the interaction between speaker embeddings and linguistically motivated phonological rules in accented speech synthesis. Using American and British English as a case study, we implement rules for flapping, rhoticity, and vowel correspondences. We propose the phoneme shift rate (PSR), a novel metric quantifying how strongly embeddings preserve or override rule-based transformations. Experiments show that combining rules with embeddings yields more authentic accents, while embeddings can attenuate or overwrite rules, revealing entanglement between accent and speaker identity. Our findings highlight rules as a lever for accent control and a framework for evaluating disentanglement in speech generation.

📄 PDF Abstract BibTeX arXiv:2601.14417

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

On the Computational Modelling of Michif Verbal Morphology

2021-04-01 · EACL 2021 2 · Fineen Davis, Eddie Antonio Santos, Heather Souter

This paper presents a finite-state computational model of the verbal morphology of Michif. Michif, the official language of the M{\'e}tis peoples, is a uniquely mixed language with Algonquian and French origins. It is sp…

Language ModelingLanguage Modelling

Learning-free L2-Accented Speech Generation using Phonological Rules

2026-03-08 · Thanathai Lertpetchpun, Yoonjeong Lee, Jihwan Lee, Tiantian Feng 외 arxiv

Accent plays a crucial role in speaker identity and inclusivity in speech technologies. Existing accented text-to-speech (TTS) systems either require large-scale accented datasets or lack fine-grained phoneme-level contr…

How Do Language Models Represent and Use Phonological Information for Allomorph Selection?

2026-09-04 · Sangwoo Kim, Sangah Lee arxiv

Language models are trained on tokenized text that obscures the sound structure of words, yet they reliably produce morphemes whose form is phonologically conditioned. It remains unclear whether they rely on item-specifi…

Cross-lingual Low Resource Speaker Adaptation Using Phonological Features

2021-11-17 · Georgia Maniati, Nikolaos Ellinas, Konstantinos Markopoulos, Georgios Vamvoukakis 외

The idea of using phonological features instead of phonemes as input to sequence-to-sequence TTS has been recently proposed for zero-shot multilingual speech synthesis. This approach is useful for code-switching, as it f…

Speech Synthesis

Training-Free Cross-Lingual Dysarthria Severity Assessment via Phonological Subspace Analysis in Self-Supervised Speech Representations

2026-04-11 · Bernard Muller, Antonio Armando Ortiz Barrañón, LaVonne Roberts arxiv

Dysarthric speech severity assessment typically requires trained clinicians or supervised models built from labelled pathological speech, limiting scalability across languages and clinical settings. We present a training…