paper-with-me

홈 › Papers

An Empirical Analysis of Deep Audio-Visual Models for Speech Recognition

2018-12-21 · Devesh Walawalkar, Yihui He, Rohit Pillai

In this project, we worked on speech recognition, specifically predicting individual words based on both the video frames and audio. Empowered by convolutional neural networks, the recent speech recognition and lip reading models are comparable to human level performance. We re-implemented and made derivations of the state-of-the-art model. Then, we conducted rich experiments including the effectiveness of attention mechanism, more accurate residual network as the backbone with pre-trained weights and the sensitivity of our model with respect to audio input with/without noise.

📄 PDF Abstract BibTeX arXiv:1812.09336

Code (0)

등록된 구현이 없습니다.

Tasks

Lip ReadingSensitivityspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Recognition of Isolated Words using Zernike and MFCC features for Audio Visual Speech Recognition

2014-07-04 · Prashant Bordea, Amarsinh Varpeb, Ramesh Manzac, Pravin Yannawara

Automatic Speech Recognition (ASR) by machine is an attractive research topic in signal processing domain and has attracted many researchers to contribute in this area. In recent year, there have been many advances in au…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+2

AV-CPL: Continuous Pseudo-Labeling for Audio-Visual Speech Recognition

2023-09-29 · Andrew Rouditchenko, Ronan Collobert, Tatiana Likhomanenko

Audio-visual speech contains synchronized audio and visual information that provides cross-modal supervision to learn representations for both automatic speech recognition (ASR) and visual speech recognition (VSR). We in…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Pseudo Label+3

On Robustness to Missing Video for Audiovisual Speech Recognition

2023-12-13 · Oscar Chang, Otavio Braga, Hank Liao, Dmitriy Serdyuk 외

It has been shown that learning audiovisual features can lead to improved speech recognition performance over audio-only features, especially for noisy speech. However, in many common applications, the visual features ar…

speech-recognitionSpeech Recognition

Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective

2024-09-29 · Chen Chen, Xiaolou Li, Zehua Liu, Lantian Li 외

In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip reading, audio-visual speech recognition, a…

Audio-Visual Speech RecognitionLip Readingspeech-recognitionSpeech Recognition+2

OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

2023-01-16 · Jeongkyun Park, Jung-Wook Hwang, Kwanghee Choi, Seung-Hyun Lee 외

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce dependencies with various prediction models d…

Audio-Visual Speech RecognitionLip ReadingSpeaker Recognitionspeech-recognition+2