paper-with-me

Papers

ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations

2024-01-01 · CVPR 2024 1 · Yuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin Chen

We propose a novel strategy ES3 for self-supervised learning of robust audio-visual speech representations from unlabeled talking face videos. While many recent approaches for this task primarily rely on guiding the learning process using the audio modality alone to capture information shared between audio and video we reframe the problem as the acquisition of shared unique (modality-specific) and synergistic speech information to address the inherent asymmetry between the modalities. Based on this formulation we propose a novel "evolving" strategy that progressively builds joint audio-visual speech representations that are strong for both uni-modal (audio & visual) and bi-modal (audio-visual) speech. First we leverage the more easily learnable audio modality to initialize audio and visual representations by capturing audio-unique and shared speech information. Next we incorporate video-unique speech information and bootstrap the audio-visual representations on top of the previously acquired shared knowledge. Finally we maximize the total audio-visual speech information including synergistic information to obtain robust and comprehensive representations. We implement ES3 as a simple Siamese framework and experiments on both English benchmarks and a newly contributed large-scale Mandarin dataset show its effectiveness. In particular on LRS2-BBC our smallest model is on par with SoTA models with only 1/2 parameters and 1/8 unlabeled data (223h).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech RecognitionLipreadingSelf-Supervised LearningSpeech Recognition

Similar Papers 제목 키워드 기반

Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision

2020-07-08 · Abhinav Shukla, Stavros Petridis, Maja Pantic

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and a…

Acoustic Scene ClassificationAction RecognitionScene ClassificationSelf-Supervised Learning+1

Jointly Learning Visual and Auditory Speech Representations from Raw Data

2022-12-12 · Alexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis 외

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets…

Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1

Does Visual Self-Supervision Improve Learning of Speech Representations for Emotion Recognition?

2020-05-04 · Abhinav Shukla, Stavros Petridis, Maja Pantic

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between au…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionFace Reconstruction+5

Visually Guided Self Supervised Learning of Speech Representations

2020-01-13 · Abhinav Shukla, Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis 외

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there …

Emotion RecognitionRepresentation LearningSelf-Supervised LearningSpeech Emotion Recognition+2

Audio-Visual Speech Enhancement and Separation by Utilizing Multi-Modal Self-Supervised Embeddings

2022-10-31 · I-Chun Chern, Kuo-Hsuan Hung, Yi-Ting Chen, Tassadaq Hussain 외

AV-HuBERT, a multi-modal self-supervised learning model, has been shown to be effective for categorical problems such as automatic speech recognition and lip-reading. This suggests that useful audio-visual speech represe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Lip Readingregression+5