MobiVSR: A Visual Speech Recognition Solution for Mobile Devices
Visual speech recognition (VSR) is the task of recognizing spoken language from video input only, without any audio. VSR has many applications as an assistive technology, especially if it could be deployed in mobile devices and embedded systems. The need of intensive computational resources and large memory footprint are two of the major obstacles in developing neural network models for VSR in a resource constrained environment. We propose a novel end-to-end deep neural network architecture for word level VSR called MobiVSR with a design parameter that aids in balancing the model's accuracy and parameter count. We use depthwise-separable 3D convolution for the first time in the domain of VSR and show how it makes our model efficient. MobiVSR achieves an accuracy of 73\% on a challenging Lip Reading in the Wild dataset with 6 times fewer parameters and 20 times lesser memory footprint than the current state of the art. MobiVSR can also be compressed to 6 MB by applying post training quantization.
Code (0)
등록된 구현이 없습니다.
Tasks
Lip ReadingQuantizationspeech-recognitionSpeech RecognitionVisual Speech RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices
Audio-visual speech recognition (AVSR) is one of the most promising solutions for reliable speech recognition, particularly when audio is corrupted by noise. Additional visual information can be used for both automatic l…
Audio-Visual Speech RecognitionGesture RecognitionLip ReadingSign Language Recognition+3Flowchase: a Mobile Application for Pronunciation Training
In this paper, we present a solution for providing personalized and instant feedback to English learners through a mobile application, called Flowchase, that is connected to a speech technology able to segment and analyz…
Representation LearningSpeech Representation LearningVulnerability of Automatic Identity Recognition to Audio-Visual Deepfakes
The task of deepfakes detection is far from being solved by speech or vision researchers. Several publicly available databases of fake synthetic video and speech were built to aid the development of detection methods. Ho…
Face RecognitionFace SwappingSpeaker Recognitiontext-to-speech+2Team HYU ASML ROBOVOX SP Cup 2024 System Description
This report describes the submission of HYU ASML team to the IEEE Signal Processing Cup 2024 (SP Cup 2024). This challenge, titled "ROBOVOX: Far-Field Speaker Recognition by a Mobile Robot," focuses on speaker recognitio…
Data AugmentationSpeaker RecognitionA Small-footprint Acoustic Echo Cancellation Solution for Mobile Full-Duplex Speech Interactions
In full-duplex speech interaction systems, effective Acoustic Echo Cancellation (AEC) is crucial for recovering echo-contaminated speech. This paper presents a neural network-based AEC solution to address challenges in m…
Speech RecognitionActivity DetectionData Augmentation