paper-with-me

홈 › Papers

Is Lip Region-of-Interest Sufficient for Lipreading?

2022-05-28 · Jing-Xuan Zhang, Gen-Shun Wan, Jia Pan

Lip region-of-interest (ROI) is conventionally used for visual input in the lipreading task. Few works have adopted the entire face as visual input because lip-excluded parts of the face are usually considered to be redundant and irrelevant to visual speech recognition. However, faces contain much more detailed information than lips, such as speakers' head pose, emotion, identity etc. We argue that such information might benefit visual speech recognition if a powerful feature extractor employing the entire face is trained. In this work, we propose to adopt the entire face for lipreading with self-supervised learning. AV-HuBERT, an audio-visual multi-modal self-supervised learning framework, was adopted in our experiments. Our experimental results showed that adopting the entire face achieved 16% relative word error rate (WER) reduction on the lipreading task, compared with the baseline method using lip as visual input. Without self-supervised pretraining, the model with face input achieved a higher WER than that using lip input in the case of limited training data (30 hours), while a slightly lower WER when using large amount of training data (433 hours).

📄 PDF Abstract BibTeX arXiv:2205.14295

Code (0)

등록된 구현이 없습니다.

Tasks

LipreadingSelf-Supervised Learningspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation

2025-02-09 · Jing-Xuan Zhang, Tingzhi Mao, Longjiang Guo, Jin Li 외

Lipreading is an important technique for facilitating human-computer interaction in noisy environments. Our previously developed self-supervised learning method, AV2vec, which leverages multimodal self-distillation, has …

Cross-Lingual TransferLipreadingSelf-Supervised LearningTransfer Learning

LCANet: End-to-End Lipreading with Cascaded Attention-CTC

2018-03-13 · Kai Xu, Dawei Li, Nick Cassimatis, Xiaolong Wang

Machine lipreading is a special type of automatic speech recognition (ASR) which transcribes human speech by visually interpreting the movement of related face regions including lips, face, and tongue. Recently, deep neu…

Cross-Attention Fusion of Visual and Geometric Features for Large Vocabulary Arabic Lipreading

2024-02-18 · Samar Daou, Achraf Ben-Hamadou, Ahmed Rekik, Abdelaziz Kallel

Lipreading involves using visual data to recognize spoken words by analyzing the movements of the lips and surrounding area. It is a hot research topic with many potential applications, such as human-machine interaction …

LipreadingLip Readingspeech-recognitionSpeech Recognition

Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions

2025-02-01 · David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

Visual speech recognition remains an open research problem where different challenges must be considered by dispensing with the auditory sense, such as visual ambiguities, the inter-personal variability among speakers, a…

Lipreadingspeech-recognitionSpeech RecognitionVisual Speech Recognition

Visual Speech Enhancement

2017-11-23 · Aviv Gabbay, Asaph Shamir, Shmuel Peleg

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved …

LipreadingSpeech Enhancement