paper-with-me

홈 › Papers

Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling

2025-09-23 · Niclas Pokel, Pehuén Moure, Roman Böhringer, Yingqiang Gao arxiv

ASR systems struggle with non-normative speech due to high acoustic variability and data scarcity. We propose a data-efficient method using phoneme-level uncertainty to guide fine-tuning for personalization. Instead of computationally expensive ensembles, we leverage Variational Low-Rank Adaptation (VI LoRA) to estimate epistemic uncertainty in foundation models. These estimates form a composite Phoneme Difficulty Score (PhDScore) that drives a targeted oversampling strategy. Evaluated on English and German datasets, including a longitudinal analysis against two clinical reports taken one year apart, we demonstrate that: (1) VI LoRA-based uncertainty aligns better with expert clinical assessments than standard entropy; (2) PhDScore captures stable, persistent articulatory difficulties; and (3) uncertainty-guided sampling significantly improves ASR accuracy for impaired speech.

📄 PDF Abstract BibTeX arXiv:2509.20396

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Demonstration of Adapt4Me: An Uncertainty-Aware Authoring Environment for Personalizing Automatic Speech Recognition to Non-normative Speech

2026-03-20 · Niclas Pokel, Yiming Zhao, Pehuén Moure, Yingqiang Gao 외 arxiv

Personalizing Automatic Speech Recognition (ASR) for non-normative speech remains challenging because data collection is labor-intensive and model training is technically complex. To address these limitations, we propose…

Speech RecognitionActive Learning

Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studies

2025-09-20 · Vishnu Raja, Adithya V Ganesan, Anand Syamkumar, Ritwik Banerjee 외 arxiv

State-of-the-art automatic speech recognition (ASR) models like Whisper, perform poorly on atypical speech, such as that produced by individuals with dysarthria. Past works for atypical speech have mostly investigated fu…

Speech Recognition

Speech Intelligibility Assessment of Dysarthric Speech by using Goodness of Pronunciation with Uncertainty Quantification

2023-05-28 · Eun Jung Yeo, Kwanghee Choi, Sunhee Kim, Minhwa Chung

This paper proposes an improved Goodness of Pronunciation (GoP) that utilizes Uncertainty Quantification (UQ) for automatic speech intelligibility assessment for dysarthric speech. Current GoP methods rely heavily on neu…

Uncertainty Quantification

Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition

2025-09-23 · Niclas Pokel, Pehuén Moure, Roman Boehringer, Shih-Chii Liu 외 arxiv

Speech impairments resulting from congenital disorders, such as cerebral palsy, down syndrome, or apert syndrome, as well as acquired brain injuries due to stroke, traumatic accidents, or tumors, present major challenges…

Speech Recognition

SE4Lip: Speech-Lip Encoder for Talking Head Synthesis to Solve Phoneme-Viseme Alignment Ambiguity

2025-04-08 · Yihuan Huang, Jiajun Liu, Yanzhen Ren, Wuyang Liu 외

Speech-driven talking head synthesis tasks commonly use general acoustic features (such as HuBERT and DeepSpeech) as guided speech features. However, we discovered that these features suffer from phoneme-viseme alignment…

3DGScross-modal alignmentNeRF