Real-time low-resource phoneme recognition on edge devices
While speech recognition has seen a surge in interest and research over the last decade, most machine learning models for speech recognition either require large training datasets or lots of storage and memory. Combined with the prominence of English as the number one language in which audio data is available, this means most other languages currently lack good speech recognition models. The method presented in this paper shows how to create and train models for speech recognition in any language which are not only highly accurate, but also require very little storage, memory and training data when compared with traditional models. This allows training models to recognize any language and deploying them on edge devices such as mobile phones or car displays for fast real-time speech recognition.
Code (1)
Tasks
Phoneme Recognitionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Optimizing Two-Pass Cross-Lingual Transfer Learning: Phoneme Recognition and Phoneme to Grapheme Translation
This research optimizes two-pass cross-lingual transfer learning in low-resource languages by enhancing phoneme recognition and phoneme-to-grapheme translation models. Our approach optimizes these two stages to improve s…
Cross-Lingual TransferPhoneme Recognitionspeech-recognitionSpeech Recognition+1Phoneme Recognition through Fine Tuning of Phonetic Representations: a Case Study on Luhya Language Varieties
Models pre-trained on multiple languages have shown significant promise for improving speech recognition, particularly for low-resource languages. In this work, we focus on phoneme recognition using Allosaurus, a method …
Phoneme Recognitionspeech-recognitionSpeech RecognitionMultilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
Labeled audio data is insufficient to build satisfying speech recognition systems for most of the languages in the world. There have been some zero-resource methods trying to perform phoneme or word-level speech recognit…
Language ModelingLanguage ModellingPhoneme Recognitionspeech-recognition+1Goodness-of-pronunciation without phoneme time alignment
In speech evaluation, an Automatic Speech Recognition (ASR) model often computes time boundaries and phoneme posteriors for input features. However, limited data for ASR training hinders expansion of speech evaluation to…
Speech RecognitionSelf-supervised Semantic-driven Phoneme Discovery for Zero-resource Speech Recognition
Phonemes are defined by their relationship to words: changing a phoneme changes the word. Learning a phoneme inventory with little supervision has been a longstanding challenge with important applications to under-resour…
Phoneme RecognitionRepresentation LearningSelf-Supervised Learningspeech-recognition+1