An Empirical Recipe for Universal Phone Recognition
Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive. Highly performant English-focused models do not generalize across languages, while multilingual models underutilize pretrained representations. It also remains unclear how data scale, architecture, and training objective contribute to multilingual PR. We present PhoneticXEUS -- trained on large-scale multilingual data and achieving state-of-the-art performance on both multilingual (17.7% PFER) and accented English speech (10.6% PFER). Through controlled ablations with evaluations across 100+ languages under a unified scheme, we empirically establish our training recipe and quantify the impact of SSL representations, data scale, and loss objectives. In addition, we analyze error patterns across language families, accented speech, and articulatory features. All data and code are released openly.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Automatic Speech Recognition and Topic Identification for Almost-Zero-Resource Languages
Automatic speech recognition (ASR) systems often need to be developed for extremely low-resource languages to serve end-uses such as audio content categorization and search. While universal phone recognition is natural t…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Humanitarianspeech-recognition+1Universal Phone Recognition for Language Agnostic Keyword Search
Recently, significant advanced have been made in universal phone recognition. Certain of these methods allow researchers to recognize phones in thousands of languages. In this paper, we explore the usage of such universa…
Differentiable Allophone Graphs for Language-Universal Speech Recognition
Building language-universal speech recognition systems entails producing phonological units of spoken sound that can be shared across languages. While speech annotations at the language-specific phoneme or surface levels…
speech-recognitionSpeech RecognitionRecipeSnap -- a lightweight image-to-recipe model
In this paper we want to address the problem of automation for recognition of photographed cooking dishes and generating the corresponding food recipes. Current image-to-recipe models are computation expensive and requir…
modelTusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition Experiments
There is growing interest in ASR systems that can recognize phones in a language-independent fashion. There is additionally interest in building language technologies for low-resource and endangered languages. However, t…