Can DNNs Learn to Lipread Full Sentences?
Finding visual features and suitable models for lipreading tasks that are more complex than a well-constrained vocabulary has proven challenging. This paper explores state-of-the-art Deep Neural Network architectures for lipreading based on a Sequence to Sequence Recurrent Neural Network. We report results for both hand-crafted and 2D/3D Convolutional Neural Network visual front-ends, online monotonic attention, and a joint Connectionist Temporal Classification-Sequence-to-Sequence loss. The system is evaluated on the publicly available TCD-TIMIT dataset, with 59 speakers and a vocabulary of over 6000 words. Results show a major improvement on a Hidden Markov Model framework. A fuller analysis of performance across visemes demonstrates that the network is not only learning the language model, but actually learning to lipread.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationLanguage ModelingLanguage ModellingLipreadingSimilar Papers 제목 키워드 기반
Towards Lipreading Sentences with Active Appearance Models
Automatic lipreading has major potential impact for speech recognition, supplementing and complementing the acoustic modality. Most attempts at lipreading have been performed on small vocabulary tasks, due to a shortfall…
Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1Visual speech recognition: aligning terminologies for better understanding
We are at an exciting time for machine lipreading. Traditional research stemmed from the adaptation of audio recognition systems. But now, the computer vision community is also participating. This joining of two previous…
Lipreadingspeech-recognitionSpeech RecognitionVisual Speech RecognitionVisual Speech Language Models
Language models (LM) are very powerful in lipreading systems. Language models built upon the ground truth utterances of datasets learn grammar and structure rules of words and sentences (the latter in the case of continu…
Language ModelingLanguage ModellingLipreadingTowards MOOCs for Lipreading: Using Synthetic Talking Heads to Train Humans in Lipreading at Scale
Many people with some form of hearing loss consider lipreading as their primary mode of day-to-day communication. However, finding resources to learn or improve one's lipreading skills can be challenging. This is further…
LipreadingLip Readingtext-to-speechText to SpeechTarget Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation
Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has …
Cross-Lingual TransferLipreadingSelf-Supervised LearningTransfer Learning