Speaker-independent machine lip-reading with speaker-dependent viseme classifiers
In machine lip-reading, which is identification of speech from visual-only information, there is evidence to show that visual speech is highly dependent upon the speaker [1]. Here, we use a phoneme-clustering method to form new phoneme-to-viseme maps for both individual and multiple speakers. We use these maps to examine how similarly speakers talk visually. We conclude that broadly speaking, speakers have the same repertoire of mouth gestures, where they differ is in the use of the gestures.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringLip ReadingSimilar Papers 제목 키워드 기반
The speaker-independent lipreading play-off; a survey of lipreading machines
Lipreading is a difficult gesture classification task. One problem in computer lipreading is speaker-independence. Speaker-independence means to achieve the same accuracy on test speakers not included in the training set…
General ClassificationLipreadingTarget Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation
Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has …
Cross-Lingual TransferLipreadingSelf-Supervised LearningTransfer LearningMKPLS: Manifold Kernel Partial Least Squares for Lipreading and Speaker Identification
Visual speech recognition is a challenging problem, due to confusion between visual speech features. The speaker identification problem is usually coupled with speech recognition. Moreover, speaker identification is impo…
LipreadingSpeaker Identificationspeech-recognitionSpeech Recognition+1Speaker-adaptive Lip Reading with User-dependent Padding
Lip reading aims to predict speech based on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements. This makes the …
Lip Readingspeech-recognitionSpeech RecognitionImproving Speaker-Independent Lipreading with Domain-Adversarial Training
We present a Lipreading system, i.e. a speech recognition system using only visual features, which uses domain-adversarial training for speaker independence. Domain-adversarial training is integrated into the optimizatio…
Lipreadingspeech-recognitionSpeech Recognition