paper-with-me

Papers

Probing self-supervised speech models for phonetic and phonemic information: a case study in aspiration

2023-06-09 · Kinan Martin, Jon Gauthier, Canaan Breiss, Roger Levy

Textless self-supervised speech models have grown in capabilities in recent years, but the nature of the linguistic information they encode has not yet been thoroughly examined. We evaluate the extent to which these models' learned representations align with basic representational distinctions made by humans, focusing on a set of phonetic (low-level) and phonemic (more abstract) contrasts instantiated in word-initial stops. We find that robust representations of both phonetic and phonemic distinctions emerge in early layers of these models' architectures, and are preserved in the principal components of deeper layer representations. Our analyses suggest two sources for this success: some can only be explained by the optimization of the models on speech data, while some can be attributed to these models' high-dimensional architectures. Our findings show that speech-trained HuBERT derives a low-noise and low-dimensional subspace corresponding to abstract phonological distinctions.

📄 PDF Abstract BibTeX arXiv:2306.06232

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Probing in the Wild: A Case Study of Self-Supervised Speech Representations on Mandarin Sub-dialects with Unsupervised Articulatory Analysis

2026-06-24 · Shu Shang, Fuliang Weng, Zeqian Hu, Yaqian Zhou arxiv

While self-supervised speech models have achieved strong performance across speech tasks, relatively little is known about how their internal phonetic representations behave under fine-grained dialect variation. Existing…

Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations

2024-06-13 · Mukhtar Mohamed, Oli Danyi Liu, Hao Tang, Sharon Goldwater

Self-supervised speech representations can hugely benefit downstream speech technologies, yet the properties that make them useful are still poorly understood. Two candidate properties related to the geometry of the repr…

Statistical word segmentation in spontaneous child-directed speech of Korean

2022-01-20 · ACL ARR January 2022 1 · Anonymous

The present study demonstrates advantages of child-directed speech (CDS) over adult-directed speech (ADS) in statistical word segmentation of spontaneous Korean. We derived phonetic input from phonemic corpus by applying…

Segmentation

ViSpeechFormer: A Phonemic Approach for Vietnamese Automatic Speech Recognition

2026-02-10 · Khoa Anh Nguyen, Long Minh Hoang, Nghia Hieu Nguyen, Luan Thanh Nguyen 외 arxiv

Vietnamese has a phonetic orthography, where each grapheme corresponds to at most one phoneme and vice versa. Exploiting this high grapheme-phoneme transparency, we propose ViSpeechFormer (\textbf{Vi}etnamese \textbf{Spe…

Speech Recognition

Self-supervised models of audio effectively explain human cortical responses to speech

2022-05-27 · Aditya R. Vaidya, Shailee Jain, Alexander G. Huth

Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on…

Representation LearningSpeech Representation Learning