An Empirical Analysis of Deep Audio-Visual Models for Speech Recognition
In this project, we worked on speech recognition, specifically predicting individual words based on both the video frames and audio. Empowered by convolutional neural networks, the recent speech recognition and lip reading models are comparable to human level performance. We re-implemented and made derivations of the state-of-the-art model. Then, we conducted rich experiments including the effectiveness of attention mechanism, more accurate residual network as the backbone with pre-trained weights and the sensitivity of our model with respect to audio input with/without noise.
Code (0)
등록된 구현이 없습니다.
Tasks
Lip ReadingSensitivityspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Recognition of Isolated Words using Zernike and MFCC features for Audio Visual Speech Recognition
Automatic Speech Recognition (ASR) by machine is an attractive research topic in signal processing domain and has attracted many researchers to contribute in this area. In recent year, there have been many advances in au…
Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+2AV-CPL: Continuous Pseudo-Labeling for Audio-Visual Speech Recognition
Audio-visual speech contains synchronized audio and visual information that provides cross-modal supervision to learn representations for both automatic speech recognition (ASR) and visual speech recognition (VSR). We in…
Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Pseudo Label+3On Robustness to Missing Video for Audiovisual Speech Recognition
It has been shown that learning audiovisual features can lead to improved speech recognition performance over audio-only features, especially for noisy speech. However, in many common applications, the visual features ar…
speech-recognitionSpeech RecognitionQuantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip reading, audio-visual speech recognition, a…
Audio-Visual Speech RecognitionLip Readingspeech-recognitionSpeech Recognition+2OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset
Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce dependencies with various prediction models d…
Audio-Visual Speech RecognitionLip ReadingSpeaker Recognitionspeech-recognition+2