SparseVSR: Lightweight and Noise Robust Visual Speech Recognition
Recent advances in deep neural networks have achieved unprecedented success in visual speech recognition. However, there remains substantial disparity between current methods and their deployment in resource-constrained devices. In this work, we explore different magnitude-based pruning techniques to generate a lightweight model that achieves higher performance than its dense model equivalent, especially under the presence of visual noise. Our sparse models achieve state-of-the-art results at 10% sparsity on the LRS3 dataset and outperform the dense equivalent up to 70% sparsity. We evaluate our 50% sparse model on 7 different visual noise types and achieve an overall absolute improvement of more than 2% WER compared to the dense equivalent. Our results confirm that sparse networks are more resistant to noise than dense networks.
Code (0)
등록된 구현이 없습니다.
Tasks
speech-recognitionSpeech RecognitionVisual Speech RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Audio Lottery: Speech Recognition Made Ultra-Lightweight, Noise-Robust, and Transferable
Lightweight speech recognition models have seen explosive demands owing to a growing amount of speech-interactive features on mobile devices. Since designing such systems from scratch is non-trivial, practitioners typica…
speech-recognitionSpeech RecognitionVisual Context-driven Audio Feature Enhancement for Robust End-to-End Audio-Visual Speech Recognition
This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy …
Audio-Visual Speech RecognitionDecoderNoisy Speech Recognitionspeech-recognition+2Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition
Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most existing efforts still focus on audio modali…
Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech RecognitionAVFormer: Injecting Vision into Frozen Speech Models for Zero-Shot AV-ASR
Audiovisual automatic speech recognition (AV-ASR) aims to improve the robustness of a speech recognition system by incorporating visual information. Training fully supervised multimodal models for this task from scratch,…
Automatic Speech RecognitionDomain AdaptationRobust Speech Recognitionspeech-recognition+1Visual-Aware Speech Recognition for Noisy Scenarios
Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Vi…
Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+2