paper-with-me

Papers

Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment

2025-02-10 · Kwanghee Choi, Eunjung Yeo, Kalvin Chang, Shinji Watanabe, David Mortensen

Allophony refers to the variation in the phonetic realization of a phoneme based on its phonetic environment. Modeling allophones is crucial for atypical pronunciation assessment, which involves distinguishing atypical from typical pronunciations. However, recent phoneme classifier-based approaches often simplify this by treating various realizations as a single phoneme, bypassing the complexity of modeling allophonic variation. Motivated by the acoustic modeling capabilities of frozen self-supervised speech model (S3M) features, we propose MixGoP, a novel approach that leverages Gaussian mixture models to model phoneme distributions with multiple subclusters. Our experiments show that MixGoP achieves state-of-the-art performance across four out of five datasets, including dysarthric and non-native speech. Our analysis further suggests that S3M features capture allophonic variation more effectively than MFCCs and Mel spectrograms, highlighting the benefits of integrating MixGoP with S3M features.

📄 PDF Abstract BibTeX arXiv:2502.07029

Code (1)

juice500ml/acoustic-units-for-ood 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Unsupervised Accent Adaptation Through Masked Language Model Correction Of Discrete Self-Supervised Speech Units

2023-09-25 · Jakob Poncelet, Hugo Van hamme

Self-supervised pre-trained speech models have strongly improved speech recognition, yet they are still sensitive to domain shifts and accented or atypical speech. Many of these models rely on quantisation or clustering …

Accented Speech RecognitionLanguage ModelingLanguage Modellingspeech-recognition+1

Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studies

2025-09-20 · Vishnu Raja, Adithya V Ganesan, Anand Syamkumar, Ritwik Banerjee 외 arxiv

State-of-the-art automatic speech recognition (ASR) models like Whisper, perform poorly on atypical speech, such as that produced by individuals with dysarthria. Past works for atypical speech have mostly investigated fu…

Speech Recognition

Don't Stop Self-Supervision: Accent Adaptation of Speech Representations via Residual Adapters

2023-07-02 · Anshu Bhatia, Sanchit Sinha, Saket Dingliwal, Karthik Gopalakrishnan 외

Speech representations learned in a self-supervised fashion from massive unlabeled speech corpora have been adapted successfully toward several downstream tasks. However, such representations may be skewed toward canonic…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Improving fairness for spoken language understanding in atypical speech with Text-to-Speech

2023-11-16 · Helin Wang, Venkatesh Ravichandran, Milind Rao, Becky Lammers 외

Spoken language understanding (SLU) systems often exhibit suboptimal performance in processing atypical speech, typically caused by neurological conditions and motor impairments. Recent advancements in Text-to-Speech (TT…

Data AugmentationFairnessSpoken Language Understandingtext-to-speech+2

Automatic Severity Classification of Dysarthric speech by using Self-supervised Model with Multi-task Learning

2022-10-27 · Eun Jung Yeo, Kwanghee Choi, Sunhee Kim, Minhwa Chung

Automatic assessment of dysarthric speech is essential for sustained treatments and rehabilitation. However, obtaining atypical speech is challenging, often leading to data scarcity issues. To tackle the problem, we prop…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClassificationMulti-Task Learning+2