The speaker-independent lipreading play-off; a survey of lipreading machines
Lipreading is a difficult gesture classification task. One problem in computer lipreading is speaker-independence. Speaker-independence means to achieve the same accuracy on test speakers not included in the training set as speakers within the training set. Current literature is limited on speaker-independent lipreading, the few independent test speaker accuracy scores are usually aggregated within dependent test speaker accuracies for an averaged performance. This leads to unclear independent results. Here we undertake a systematic survey of experiments with the TCD-TIMIT dataset using both conventional approaches and deep learning methods to provide a series of wholly speaker-independent benchmarks and show that the best speaker-independent machine scores 69.58% accuracy with CNN features and an SVM classifier. This is less than state of the art speaker-dependent lipreading machines, but greater than previously reported in independence experiments.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationLipreadingSimilar Papers 제목 키워드 기반
Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation
Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has …
Cross-Lingual TransferLipreadingSelf-Supervised LearningTransfer LearningLipper: Synthesizing Thy Speech using Multi-View Lipreading
Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos int…
LipreadingMKPLS: Manifold Kernel Partial Least Squares for Lipreading and Speaker Identification
Visual speech recognition is a challenging problem, due to confusion between visual speech features. The speaker identification problem is usually coupled with speech recognition. Moreover, speaker identification is impo…
LipreadingSpeaker Identificationspeech-recognitionSpeech Recognition+1Learning Speaker-Invariant Visual Features for Lipreading
Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract visual features that include speaker-specific lip attributes (e.g., shape…
DisentanglementLipreadingSpeaker RecognitionTowards MOOCs for Lipreading: Using Synthetic Talking Heads to Train Humans in Lipreading at Scale
Many people with some form of hearing loss consider lipreading as their primary mode of day-to-day communication. However, finding resources to learn or improve one's lipreading skills can be challenging. This is further…
LipreadingLip Readingtext-to-speechText to Speech