paper-with-me

홈 › Papers

Deep Learning for Lip Reading using Audio-Visual Information for Urdu Language

2018-02-15 · M Faisal, Sanaullah Manzoor

Human lip-reading is a challenging task. It requires not only knowledge of underlying language but also visual clues to predict spoken words. Experts need certain level of experience and understanding of visual expressions learning to decode spoken words. Now-a-days, with the help of deep learning it is possible to translate lip sequences into meaningful words. The speech recognition in the noisy environments can be increased with the visual information [1]. To demonstrate this, in this project, we have tried to train two different deep-learning models for lip-reading: first one for video sequences using spatiotemporal convolution neural network, Bi-gated recurrent neural network and Connectionist Temporal Classification Loss, and second for audio that inputs the MFCC features to a layer of LSTM cells and output the sequence. We have also collected a small audio-visual dataset to train and test our model. Our target is to integrate our both models to improve the speech recognition in the noisy environment

📄 PDF Abstract BibTeX arXiv:1802.05521

Code (0)

등록된 구현이 없습니다.

Tasks

Lip Readingspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

UQuAD1.0: Development of an Urdu Question Answering Training Data for Machine Reading Comprehension

2021-11-02 · Samreen Kazi, Shakeel Khoja

In recent years, low-resource Machine Reading Comprehension (MRC) has made significant progress, with models getting remarkable performance on various language datasets. However, none of these models have been customized…

ArticlesMachine Reading ComprehensionMachine TranslationQuestion Answering+1

Can You Hear, Localize, and Segment Continually? An Exemplar-Free Continual Learning Benchmark for Audio-Visual Segmentation

2026-03-09 · Siddeshwar Raghavan, Gautham Vinod, Bruce Coburn, Fengqing Zhu arxiv

Audio-Visual Segmentation (AVS) aims to produce pixel-level masks of sound producing objects in videos, by jointly learning from audio and visual signals. However, real-world environments are inherently dynamic, causing …

Continual Learning

Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading

2022-04-04 · The AAAI Conference on Artificial Intelligence (AAAI) 2022 3 · Minsu Kim, Jeong Hun Yeo, Yong Man Ro

Recognizing speech from silent lip movement, which is called lip reading, is a challenging task due to 1) the inherent information insufficiency of lip movement to fully represent the speech, and 2) the existence of homo…

LipreadingLip Reading

Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective

2024-09-29 · Chen Chen, Xiaolou Li, Zehua Liu, Lantian Li 외

In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip reading, audio-visual speech recognition, a…

Audio-Visual Speech RecognitionLip Readingspeech-recognitionSpeech Recognition+2

A Benchmark Dataset and a Framework for Urdu Multimodal Named Entity Recognition

2025-05-08 · Hussain Ahmad, Qingyang Zeng, Jing Wan

The emergence of multimodal content, particularly text and images on social media, has positioned Multimodal Named Entity Recognition (MNER) as an increasingly important area of research within Natural Language Processin…

named-entity-recognitionNamed Entity Recognition