paper-with-me

Papers

AV Taris: Online Audio-Visual Speech Recognition

2020-12-14 · George Sterpu, Naomi Harte

In recent years, Automatic Speech Recognition (ASR) technology has approached human-level performance on conversational speech under relatively clean listening conditions. In more demanding situations involving distant microphones, overlapped speech, background noise, or natural dialogue structures, the ASR error rate is at least an order of magnitude higher. The visual modality of speech carries the potential to partially overcome these challenges and contribute to the sub-tasks of speaker diarisation, voice activity detection, and the recovery of the place of articulation, and can compensate for up to 15dB of noise on average. This article develops AV Taris, a fully differentiable neural network model capable of decoding audio-visual speech in real time. We achieve this by connecting two recently proposed models for audio-visual speech integration and online speech recognition, namely AV Align and Taris. We evaluate AV Taris under the same conditions as AV Align and Taris on one of the largest publicly available audio-visual speech datasets, LRS2. Our results show that AV Taris is superior to the audio-only variant of Taris, demonstrating the utility of the visual modality to speech recognition within the real time decoding framework defined by Taris. Compared to an equivalent Transformer-based AV Align model that takes advantage of full sentences without meeting the real-time requirement, we report an absolute degradation of approximately 3% with AV Taris. As opposed to the more popular alternative for online speech recognition, namely the RNN Transducer, Taris offers a greatly simplified fully differentiable training pipeline. As a consequence, AV Taris has the potential to popularise the adoption of Audio-Visual Speech Recognition (AVSR) technology and overcome the inherent limitations of the audio modality in less optimal listening conditions.

📄 PDF Abstract BibTeX arXiv:2012.07467

Code (1)

georgesterpu/Taris 공식 구현 tf

Tasks

Action DetectionActivity DetectionAudio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

Learning to Count Words in Fluent Speech enables Online Speech Recognition

2020-06-08 · George Sterpu, Christian Saam, Naomi Harte

Sequence to Sequence models, in particular the Transformer, achieve state of the art results in Automatic Speech Recognition. Practical usage is however limited to cases where full utterance latency is acceptable. In thi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

AV-CPL: Continuous Pseudo-Labeling for Audio-Visual Speech Recognition

2023-09-29 · Andrew Rouditchenko, Ronan Collobert, Tatiana Likhomanenko

Audio-visual speech contains synchronized audio and visual information that provides cross-modal supervision to learn representations for both automatic speech recognition (ASR) and visual speech recognition (VSR). We in…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Pseudo Label+3

MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

2023-03-01 · Mohamed Anwar, Bowen Shi, Vedanuj Goswami, Wei-Ning Hsu 외

We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languages. It is fully transcribed and covers 6…

Audio-Visual Speech RecognitionRobust Speech Recognitionspeech-recognitionSpeech Recognition+4

Visual Context-driven Audio Feature Enhancement for Robust End-to-End Audio-Visual Speech Recognition

2022-07-13 · Joanna Hong, Minsu Kim, Daehun Yoo, Yong Man Ro

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy …

Audio-Visual Speech RecognitionDecoderNoisy Speech Recognitionspeech-recognition+2

Learning Contextually Fused Audio-visual Representations for Audio-visual Speech Recognition

2022-02-15 · Zi-Qiang Zhang, Jie Zhang, Jian-Shu Zhang, Ming-Hui Wu 외

With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech re…

Audio-Visual Speech RecognitionLipreadingRepresentation LearningSelf-Supervised Learning+3