MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
Pre-trained models have been a foundational approach in speech recognition, albeit with associated additional costs. In this study, we propose a regularization technique that facilitates the training of visual and audio-visual speech recognition models (VSR and AVSR) from scratch. This approach, abbreviated as \textbf{MSRS} (Multimodal Speech Recognition from Scratch), introduces a sparse regularization that rapidly learns sparse structures within the dense model at the very beginning of training, which receives healthier gradient flow than the dense equivalent. Once the sparse mask stabilizes, our method allows transitioning to a dense model or keeping a sparse model by updating non-zero values. MSRS achieves competitive results in VSR and AVSR with 21.1% and 0.9% WER on the LRS3 benchmark, while reducing training time by at least 2x. We explore other sparse approaches and show that only MSRS enables training from scratch by implicitly masking the weights affected by vanishing gradients.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech RecognitionSimilar Papers 제목 키워드 기반
Improving Unimodal Inference with Multimodal Transformers
This paper proposes an approach for improving performance of unimodal models with multimodal training. Our approach involves a multi-branch architecture that incorporates unimodal models with a multimodal transformer-bas…
Emotion RecognitionGesture RecognitionHand Gesture RecognitionHand-Gesture Recognition+1Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
Emotion recognition and sentiment analysis are pivotal tasks in speech and language processing, particularly in real-world scenarios involving multi-party, conversational data. This paper presents a multimodal approach t…
Emotion RecognitionMultimodal Emotion RecognitionSentiment AnalysisAVFormer: Injecting Vision into Frozen Speech Models for Zero-Shot AV-ASR
Audiovisual automatic speech recognition (AV-ASR) aims to improve the robustness of a speech recognition system by incorporating visual information. Training fully supervised multimodal models for this task from scratch,…
Automatic Speech RecognitionDomain AdaptationRobust Speech Recognitionspeech-recognition+1MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
This technical report describes the MERaLiON-SpeechEncoder, a foundation model designed to support a wide range of downstream speech applications. Developed as part of Singapore's National Multimodal Large Language Model…
Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3A Continual Learning Framework for Adaptive Control of Modular Soft Robots
Soft robots have attracted significant attention in applications such as medical intervention, rehabilitation, and robotic manipulation due to their inherent compliance, flexibility, and high degrees of freedom. Modular …
Continual Learning