LipFormer: Learning to Lipread Unseen Speakers based on Visual-Landmark Transformers
Lipreading refers to understanding and further translating the speech of a speaker in the video into natural language. State-of-the-art lipreading methods excel in interpreting overlap speakers, i.e., speakers appear in both training and inference sets. However, generalizing these methods to unseen speakers incurs catastrophic performance degradation due to the limited number of speakers in training bank and the evident visual variations caused by the shape/color of lips for different speakers. Therefore, merely depending on the visible changes of lips tends to cause model overfitting. To address this problem, we propose to use multi-modal features across visual and landmarks, which can describe the lip motion irrespective to the speaker identities. Then, we develop a sentence-level lipreading framework based on visual-landmark transformers, namely LipFormer. Specifically, LipFormer consists of a lip motion stream, a facial landmark stream, and a cross-modal fusion. The embeddings from the two streams are produced by self-attention, which are fed to the cross-attention module to achieve the alignment between visuals and landmarks. Finally, the resulting fused features can be decoded to output texts by a cascade seq2seq model. Experiments demonstrate that our method can effectively enhance the model generalization to unseen speakers.
Code (0)
등록된 구현이 없습니다.
Tasks
LipreadingSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Cross-Attention Fusion of Visual and Geometric Features for Large Vocabulary Arabic Lipreading
Lipreading involves using visual data to recognize spoken words by analyzing the movements of the lips and surrounding area. It is a hot research topic with many potential applications, such as human-machine interaction …
LipreadingLip Readingspeech-recognitionSpeech RecognitionPart-based Lipreading for Audio-Visual Speech Recognition
Lipreading is an important component of audio-visual speech recognition. However, lips are usually modeled as a whole in lipreading, which ignores that each part of lip focuses on different characteristics of mouth and t…
Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation
Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has …
Cross-Lingual TransferLipreadingSelf-Supervised LearningTransfer LearningLearning Speaker-Invariant Visual Features for Lipreading
Lipreading is a challenging cross-modal task that aims to convert visual lip movements into spoken text. Existing lipreading methods often extract visual features that include speaker-specific lip attributes (e.g., shape…
DisentanglementLipreadingSpeaker RecognitionCan DNNs Learn to Lipread Full Sentences?
Finding visual features and suitable models for lipreading tasks that are more complex than a well-constrained vocabulary has proven challenging. This paper explores state-of-the-art Deep Neural Network architectures for…
General ClassificationLanguage ModelingLanguage ModellingLipreading