paper-with-me

홈 › Papers

On the Importance of Video Action Recognition for Visual Lipreading

2019-03-22 · Xinshuo Weng

We focus on the word-level visual lipreading, which requires to decode the word from the speaker's video. Recently, many state-of-the-art visual lipreading methods explore the end-to-end trainable deep models, involving the use of 2D convolutional networks (e.g., ResNet) as the front-end visual feature extractor and the sequential model (e.g., Bi-LSTM or Bi-GRU) as the back-end. Although a deep 2D convolution neural network can provide informative image-based features, it ignores the temporal motion existing between the adjacent frames. In this work, we investigate the spatial-temporal capacity power of I3D (Inflated 3D ConvNet) for visual lipreading. We demonstrate that, after pre-trained on the large-scale video action recognition dataset (e.g., Kinetics), our models show a considerable improvement of performance on the task of lipreading. A comparison between a set of video model architectures and input data representation is also reported. Our extensive experiments on LRW shows that a two-stream I3D model with RGB video and optical flow as the inputs achieves the state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:1903.09616

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionLipreadingOptical Flow EstimationTemporal Action Localization

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Cross-Attention Fusion of Visual and Geometric Features for Large Vocabulary Arabic Lipreading

2024-02-18 · Samar Daou, Achraf Ben-Hamadou, Ahmed Rekik, Abdelaziz Kallel

Lipreading involves using visual data to recognize spoken words by analyzing the movements of the lips and surrounding area. It is a hot research topic with many potential applications, such as human-machine interaction …

LipreadingLip Readingspeech-recognitionSpeech Recognition

Large-Scale Visual Speech Recognition

2018-07-13 · ICLR 2019 5 · Brendan Shillingford, Yannis Assael, Matthew W. Hoffman, Thomas Paine 외

This work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consisting of pairs of text and video clips of …

DecoderLipreadingspeech-recognitionSpeech Recognition+1

MA-LipNet: Multi-Dimensional Attention Networks for Robust Lipreading

2026-01-27 · Matteo Rossi arxiv

Lipreading, the technology of decoding spoken content from silent videos of lip movements, holds significant application value in fields such as public security. However, due to the subtle nature of articulatory gestures…

Visual Speech Recognition

LRWR: Large-Scale Benchmark for Lip Reading in Russian language

2021-09-14 · Evgeniy Egorov, Vasily Kostyumov, Mikhail Konyk, Sergey Kolesnikov

Lipreading, also known as visual speech recognition, aims to identify the speech content from videos by analyzing the visual deformations of lips and nearby areas. One of the significant obstacles for research in this fi…

LipreadingLip Readingspeech-recognitionSpeech Recognition+1

LipNet: End-to-End Sentence-level Lipreading

2016-11-05 · Yannis M. Assael, Brendan Shillingford, Shimon Whiteson, Nando de Freitas

Lipreading is the task of decoding text from the movement of a speaker's mouth. Traditional approaches separated the problem into two stages: designing or learning visual features, and prediction. More recent deep liprea…

General ClassificationLipreadingSentence