A bandit approach to curriculum generation for automatic speech recognition
The Automated Speech Recognition (ASR) task has been a challenging domain especially for low data scenarios with few audio examples. This is the main problem in training ASR systems on the data from low-resource or marginalized languages. In this paper we present an approach to mitigate the lack of training data by employing Automated Curriculum Learning in combination with an adversarial bandit approach inspired by Reinforcement learning. The goal of the approach is to optimize the training sequence of mini-batches ranked by the level of difficulty and compare the ASR performance metrics against the random training sequence and discrete curriculum. We test our approach on a truly low-resource language and show that the bandit framework has a good improvement over the baseline transfer-learning model.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)reinforcement-learningReinforcement Learning (RL)speech-recognitionSpeech RecognitionTransfer LearningSimilar Papers 제목 키워드 기반
LLMs-Integrated Automatic Hate Speech Recognition Using Controllable Text Generation Models
This paper proposes an automatic speech recognition (ASR) model for hate speech using large language models (LLMs). The proposed method integrates the encoder of the ASR model with the decoder of the LLMs, enabling simul…
Text ClassificationSpeech RecognitionText GenerationA Curriculum Learning Method for Improved Noise Robustness in Automatic Speech Recognition
The performance of automatic speech recognition systems under noisy environments still leaves room for improvement. Speech enhancement or feature enhancement techniques for increasing noise robustness of these systems us…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1DORB: Dynamically Optimizing Multiple Rewards with Bandits
Policy gradients-based reinforcement learning has proven to be a promising approach for directly optimizing non-differentiable evaluation metrics for language generation tasks. However, optimizing for a specific metric r…
Data-to-Text GenerationQuestion GenerationQuestion-GenerationText GenerationKinSPEAK: Improving speech recognition for Kinyarwanda via semi-supervised learning methods
Despite recent availability of large transcribed Kinyarwanda speech data, achieving robust speech recognition for Kinyarwanda is still challenging. In this work, we show that using self-supervised pre-training, following…
Robust Speech Recognitionspeech-recognitionSpeech RecognitionCurriculum Learning for Speech Emotion Recognition from Crowdsourced Labels
This study introduces a method to design a curriculum for machine-learning to maximize the efficiency during the training process of deep neural networks (DNNs) for speech emotion recognition. Previous studies in other m…
Emotion RecognitionMulti-class ClassificationSpeech Emotion Recognition