Deep Learning-based Spatio Temporal Facial Feature Visual Speech Recognition
In low-resource computing contexts, such as smartphones and other tiny devices, Both deep learning and machine learning are being used in a lot of identification systems. as authentication techniques. The transparent, contactless, and non-invasive nature of these face recognition technologies driven by AI has led to their meteoric rise in popularity in recent years. While they are mostly successful, there are still methods to get inside without permission by utilising things like pictures, masks, glasses, etc. In this research, we present an alternate authentication process that makes use of both facial recognition and the individual's distinctive temporal facial feature motions while they speak a password. Because the suggested methodology allows for a password to be specified in any language, it is not limited by language. The suggested model attained an accuracy of 96.1% when tested on the industry-standard MIRACL-VC1 dataset, demonstrating its efficacy as a reliable and powerful solution. In addition to being data-efficient, the suggested technique shows promising outcomes with as little as 10 positive video examples for training the model. The effectiveness of the network's training is further proved via comparisons with other combined facial recognition and lip reading models.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningFace RecognitionLip Readingspeech-recognitionSpeech RecognitionVisual Speech RecognitionSimilar Papers 제목 키워드 기반
Landmark Guided Visual Feature Extractor for Visual Speech Recognition with Limited Resource
Visual speech recognition is a technique to identify spoken content in silent speech videos, which has raised significant attention in recent years. Advancements in data-driven deep learning methods have significantly im…
Visual Speech RecognitionSpatio-Temporal Attention Mechanism and Knowledge Distillation for Lip Reading
Despite the advancement in the domain of audio and audio-visual speech recognition, visual speech recognition systems are still quite under-explored due to the visual ambiguity of some phonemes. In this work, we propose …
Audio-Visual Speech RecognitionKnowledge DistillationLip Readingspeech-recognition+2Audio-visual video face hallucination with frequency supervision and cross modality support by speech based lip reading loss
Recently, there has been numerous breakthroughs in face hallucination tasks. However, the task remains rather challenging in videos in comparison to the images due to inherent consistency issues. The presence of extra te…
Face HallucinationGenerative Adversarial NetworkHallucinationLip ReadingMSSTNet: A Multi-Scale Spatio-Temporal CNN-Transformer Network for Dynamic Facial Expression Recognition
Unlike typical video action recognition, Dynamic Facial Expression Recognition (DFER) does not involve distinct moving targets but relies on localized changes in facial muscles. Addressing this distinctive attribute, we …
Action RecognitionAttributeDynamic Facial Expression RecognitionFacial Expression Recognition+1Spatio-Temporal Transformer for Dynamic Facial Expression Recognition in the Wild
Previous methods for dynamic facial expression in the wild are mainly based on Convolutional Neural Networks (CNNs), whose local operations ignore the long-range dependencies in videos. To solve this problem, we propose …
Dynamic Facial Expression RecognitionFacial Expression RecognitionFacial Expression Recognition (FER)