paper-with-me

홈 › Papers

Scaling Human and G2P Supervision for Robust Phonetic Transcription

2026-06-14 · Alexander Metzger, Aruna Srivastava, Ruslan Mukhamedvaleev arxiv

Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech. A common alternative is using Grapheme-to-Phoneme (G2P) models to auto-generate phonetic labels from text transcripts at scale. We study how automatic phonetic transcription performance scales with human and G2P supervision in English. Using a curated 80-hour benchmark spanning native, non-native and post-stroke speech, we identify a supervision quality threshold: G2P supervision helps only when fewer than 20-30 hours of human annotation are available. Beyond this threshold, it provides no significant benefit and can reduce cross-dialect robustness. What is effective after this threshold is ASR pretraining which we use to achieve a 2.3x reduction in weighted phone feature error rate over prior systems, with strong gains on non-native and aphasic speech. These results suggest that quantity-driven G2P scaling may yield diminishing returns for robust generalization.

📄 PDF Abstract BibTeX arXiv:2606.16019

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision

2024-06-04 · Saierdaer Yusuyin, Te Ma, Hao Huang, Wenbo Zhao 외

There exist three approaches for multilingual and crosslingual automatic speech recognition (MCL-ASR) - supervised pretraining with phonetic or graphemic transcription, and self-supervised pretraining. We find that pretr…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Neural Representations for Modeling Variation in Speech

2020-11-25 · Martijn Bartelds, Wietse de Vries, Faraz Sanal, Caitlin Richter 외

Variation in speech is often quantified by comparing phonetic transcriptions of the same utterance. However, manually transcribing speech is time-consuming and error prone. As an alternative, therefore, we investigate th…

Impact of Phonetics on Speaker Identity in Adversarial Voice Attack

2025-09-18 · Daniyal Kabir Dar, Qiben Yan, Li Xiao, Arun Ross arxiv

Adversarial perturbations in speech pose a serious threat to automatic speech recognition (ASR) and speaker verification by introducing subtle waveform modifications that remain imperceptible to humans but can significan…

Speaker VerificationSpeaker RecognitionSpeech Recognition

Phonetic Segmentation of the UCLA Phonetics Lab Archive

2024-03-28 · Eleanor Chodroff, Blaž Pažon, Annie Baker, Steven Moran

Research in speech technologies and comparative linguistics depends on access to diverse and accessible speech data. The UCLA Phonetics Lab Archive is one of the earliest multilingual speech corpora, with long-form audio…

Universal Automatic Phonetic Transcription into the International Phonetic Alphabet

2023-08-07 · Chihiro Taguchi, Yusuke Sakai, Parisa Haghani, David Chiang

This paper presents a state-of-the-art model for transcribing speech in any language into the International Phonetic Alphabet (IPA). Transcription of spoken languages into IPA is an essential yet time-consuming process i…