paper-with-me

Papers

BRAVEn: Improving Self-Supervised Pre-training for Visual and Auditory Speech Recognition

2024-04-02 · Alexandros Haliassos, Andreas Zinonos, Rodrigo Mira, Stavros Petridis, Maja Pantic

Self-supervision has recently shown great promise for learning visual and auditory speech representations from unlabelled data. In this work, we propose BRAVEn, an extension to the recent RAVEn method, which learns speech representations entirely from raw audio-visual data. Our modifications to RAVEn enable BRAVEn to achieve state-of-the-art results among self-supervised methods in various settings. Moreover, we observe favourable scaling behaviour by increasing the amount of unlabelled data well beyond other self-supervised works. In particular, we achieve 20.0% / 1.7% word error rate for VSR / ASR on the LRS3 test set, with only 30 hours of labelled data and no external ASR models. Our results suggest that readily available unlabelled audio-visual data can largely replace costly transcribed data.

📄 PDF Abstract BibTeX arXiv:2404.02098

Code (1)

ahaliassos/raven 공식 구현 pytorch

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Jointly Learning Visual and Auditory Speech Representations from Raw Data

2022-12-12 · Alexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis 외

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets…

Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1

CLAR: Contrastive Learning of Auditory Representations

2020-10-19 · Haider Al-Tahan, Yalda Mohsenzadeh

Learning rich visual representations using contrastive self-supervised learning has been extremely successful. However, it is still a major question whether we could use a similar approach to learn superior auditory repr…

Contrastive LearningSelf-Supervised Learning

Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs

2024-11-04 · Alexandros Haliassos, Rodrigo Mira, Honglie Chen, Zoe Landgraf 외

Research in auditory, visual, and audiovisual speech recognition (ASR, VSR, and AVSR, respectively) has traditionally been conducted independently. Even recent self-supervised studies addressing two or all three tasks si…

Lipreadingspeech-recognitionSpeech Recognition

Show from Tell: Audio-Visual Modelling in Clinical Settings

2023-10-25 · Jianbo Jiao, Mohammad Alsharid, Lior Drukker, Aris T. Papageorghiou 외

Auditory and visual signals usually present together and correlate with each other, not only in natural environments but also in clinical settings. However, the audio-visual modelling in the latter case can be more chall…

Self-Supervised Learning

Dual Domain-Adversarial Learning for Audio-Visual Saliency Prediction

2022-08-10 · Yingzi Fan, Longfei Han, Yue Zhang, Lechao Cheng 외

Both visual and auditory information are valuable to determine the salient regions in videos. Deep convolution neural networks (CNN) showcase strong capacity in coping with the audio-visual saliency prediction task. Due …

Domain AdaptationPredictionSaliency PredictionUnsupervised Domain Adaptation