paper-with-me

홈 › Papers

Towards Lipreading Sentences with Active Appearance Models

2018-05-29

Automatic lipreading has major potential impact for speech recognition, supplementing and complementing the acoustic modality. Most attempts at lipreading have been performed on small vocabulary tasks, due to a shortfall of appropriate audio-visual datasets. In this work we use the publicly available TCD-TIMIT database, designed for large vocabulary continuous audio-visual speech recognition. We compare the viseme recognition performance of the most widely used features for lipreading, Discrete Cosine Transform (DCT) and Active Appearance Models (AAM), in a traditional Hidden Markov Model (HMM) framework. We also exploit recent advances in AAM fitting. We found the DCT to outperform AAM by more than 6% for a viseme recognition task with 56 speakers. The overall accuracy of the DCT is quite low (32-34%). We conclude that a fundamental rethink of the modelling of visual features may be needed for this task.

📄 PDF Abstract BibTeX arXiv:1805.11688

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

Visual speech recognition: aligning terminologies for better understanding

2017-10-03 · Helen L. Bear, Sarah Taylor

We are at an exciting time for machine lipreading. Traditional research stemmed from the adaptation of audio recognition systems. But now, the computer vision community is also participating. This joining of two previous…

Lipreadingspeech-recognitionSpeech RecognitionVisual Speech Recognition

Visual Speech Language Models

2018-09-14

Language models (LM) are very powerful in lipreading systems. Language models built upon the ground truth utterances of datasets learn grammar and structure rules of words and sentences (the latter in the case of continu…

Language ModelingLanguage ModellingLipreading

Can DNNs Learn to Lipread Full Sentences?

2018-05-29 · George Sterpu, Christian Saam, Naomi Harte

Finding visual features and suitable models for lipreading tasks that are more complex than a well-constrained vocabulary has proven challenging. This paper explores state-of-the-art Deep Neural Network architectures for…

General ClassificationLanguage ModelingLanguage ModellingLipreading

Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation

2025-02-09 · Jing-Xuan Zhang, Tingzhi Mao, Longjiang Guo, Jin Li 외

Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has …

Cross-Lingual TransferLipreadingSelf-Supervised LearningTransfer Learning

The speaker-independent lipreading play-off; a survey of lipreading machines

2018-10-24 · Jake Burton, David Frank, Madhi Saleh, Nassir Navab 외

Lipreading is a difficult gesture classification task. One problem in computer lipreading is speaker-independence. Speaker-independence means to achieve the same accuracy on test speakers not included in the training set…

General ClassificationLipreading