paper-with-me

홈 › Papers

Part-based Lipreading for Audio-Visual Speech Recognition

2020-12-14 · IEEE International Conference on Systems, Man, and Cybernetics (SMC) 2020 12 · Ziling Miao, Hong Liu, Bing Yang

Lipreading is an important component of audio-visual speech recognition. However, lips are usually modeled as a whole in lipreading, which ignores that each part of lip focuses on different characteristics of mouth and the overall model can not fit each part perfectly. Besides, features based on the whole lip usually vary a lot according to different speakers, which leads that the training databases usually need to contain as much speakers as possible. In this paper, A part-based lipreading (PBL) method is proposed to deal with the mismatch between an overall lip model and the separate parts of lips, also the excessive dependence of models on the speakers in training set. PBL models lips partly and predicts jointly. It employs a uniform partition strategy on convolutional features and generates several part-level sub-results for final prediction. Experiments are performed on a large publicly available dataset (LRW) and part of it (p-LRW, 65 words), in order to simulate the progressive instructions in the working scene of robots. Word accuracy of PBL reaches 82.8% on LRW and 88.9% on p-LRW. Finally, an end-to-end audio-visual speech recognition system using PBL is established and achieves 98.3% word accuracy on LRW.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

Towards Lipreading Sentences with Active Appearance Models

2018-05-29

Automatic lipreading has major potential impact for speech recognition, supplementing and complementing the acoustic modality. Most attempts at lipreading have been performed on small vocabulary tasks, due to a shortfall…

Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1

Pushing the boundaries of audiovisual word recognition using Residual Networks and LSTMs

2018-11-03 · Themos Stafylakis, Muhammad Haris Khan, Georgios Tzimiropoulos

Visual and audiovisual speech recognition are witnessing a renaissance which is largely due to the advent of deep learning methods. In this paper, we present a deep learning architecture for lipreading and audiovisual wo…

Lipreadingspeech-recognitionSpeech Recognition

Lip-Listening: Mixing Senses to Understand Lips using Cross Modality Knowledge Distillation for Word-Based Models

2022-06-05 · Hadeel Mabrouk, Omar Abugabal, Nourhan Sakr, Hesham M. Eraqi

In this work, we propose a technique to transfer speech recognition capabilities from audio speech recognition systems to visual speech recognizers, where our goal is to utilize audio data during lipreading model trainin…

Knowledge DistillationLipreadingLip Readingspeech-recognition+2

Is Lip Region-of-Interest Sufficient for Lipreading?

2022-05-28 · Jing-Xuan Zhang, Gen-Shun Wan, Jia Pan

Lip region-of-interest (ROI) is conventionally used for visual input in the lipreading task. Few works have adopted the entire face as visual input because lip-excluded parts of the face are usually considered to be redu…

LipreadingSelf-Supervised Learningspeech-recognitionSpeech Recognition+1

Visual speech recognition: aligning terminologies for better understanding

2017-10-03 · Helen L. Bear, Sarah Taylor

We are at an exciting time for machine lipreading. Traditional research stemmed from the adaptation of audio recognition systems. But now, the computer vision community is also participating. This joining of two previous…

Lipreadingspeech-recognitionSpeech RecognitionVisual Speech Recognition