paper-with-me

홈 › Papers

Lexical Tone is Hard to Quantize: Probing Discrete Speech Units in Mandarin and Yorùbá

2026-04-08 · Opeyemi Osakuade, Simon King arxiv

Discrete speech units (DSUs) are derived by quantising representations from models trained using self-supervised learning (SSL). They are a popular representation for a wide variety of spoken language tasks, including those where prosody matters. DSUs are especially convenient for tasks where text and speech are jointly modelled, such as text-to-speech and multimodal dialogue systems. But we have found that DSUs encode suprasegmental information less reliably than segmental structure, which we demonstrate in this work using lexical tone, though this limitation likely extends to other suprasegmental features such as prosody. Our investigations using the tone languages Mandarin and Yorùbá show that the SSL latent representations themselves do encode tone, yet DSUs obtained using quantisation tend to prioritise phonetic structure, which makes lexical tone less reliably encoded. This remains true for a variety of quantisation methods, not only the most common, K-means. We conclude that current DSU quantisation strategies have limitations for suprasegmental features, which suggests a need for new, tone-aware (or prosody-aware) techniques in speech representation learning. We point towards a potential form of the solution by performing K-means clustering once to encode phonetic information, then again on the residual representation, which better encodes lexical tone.

📄 PDF Abstract BibTeX arXiv:2604.07467

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningRepresentation Learning

Similar Papers 제목 키워드 기반

Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models

2025-06-12 · Michele Gubian, Ioana Krehan, Oli Liu, James Kirby 외

Analyses of self-supervised speech models have begun to reveal where and how they represent different types of information. However, almost all analyses have focused on English. Here, we examine how wav2vec2 models train…

Quantization Robustness of Monotone Operator Equilibrium Networks

2026-03-11 · James Li, Philip H. W. Leong, Thomas Chaffey arxiv

Monotone operator equilibrium networks are implicit-layer models whose output is the unique equilibrium of a monotone operator, guaranteeing existence, uniqueness, and convergence. When deployed on low-precision hardware…

A layer-wise analysis of Mandarin and English suprasegmentals in SSL speech models

2024-08-24 · Antón de la Fuente, Dan Jurafsky

This study asks how self-supervised speech models represent suprasegmental categories like Mandarin lexical tone, English lexical stress, and English phrasal accents. Through a series of probing tasks, we make layer-wise…

Specificity

Probing the phonetic and phonological knowledge of tones in Mandarin TTS models

2019-12-23 · Jian Zhu

This study probes the phonetic and phonological knowledge of lexical tones in TTS models through two experiments. Controlled stimuli for testing tonal coarticulation and tone sandhi in Mandarin were fed into Tacotron 2 a…

Communication-Channel Optimized Partition

2020-01-06 · Thuan Nguyen, Thinh Nguyen

Given an original discrete source X with the distribution p_X that is corrupted by noise to produce the noisy data Y with the given joint distribution p(X, Y). A quantizer/classifier Q : Y -> Z is then used to classify/q…