paper-with-me

Papers

Does Visual Self-Supervision Improve Learning of Speech Representations for Emotion Recognition?

2020-05-04 · Abhinav Shukla, Stavros Petridis, Maja Pantic

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between audio and visual modalities for cross-modal self-supervision. This work (1) investigates visual self-supervision via face reconstruction to guide the learning of audio representations; (2) proposes an audio-only self-supervision approach for speech representation learning; (3) shows that a multi-task combination of the proposed visual and audio self-supervision is beneficial for learning richer features that are more robust in noisy conditions; (4) shows that self-supervised pretraining can outperform fully supervised training and is especially useful to prevent overfitting on smaller sized datasets. We evaluate our learned audio representations for discrete emotion recognition, continuous affect recognition and automatic speech recognition. We outperform existing self-supervised methods for all tested downstream tasks. Our results demonstrate the potential of visual self-supervision for audio feature learning and suggest that joint visual and audio self-supervision leads to more informative audio representations for speech and emotion recognition.

📄 PDF Abstract BibTeX arXiv:2005.01400

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionFace ReconstructionRepresentation LearningSelf-Supervised Learningspeech-recognitionSpeech RecognitionSpeech Representation Learning

Similar Papers 제목 키워드 기반

Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision

2020-07-08 · Abhinav Shukla, Stavros Petridis, Maja Pantic

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and a…

Acoustic Scene ClassificationAction RecognitionScene ClassificationSelf-Supervised Learning+1

Large-Scale Self- and Semi-Supervised Learning for Speech Translation

2021-04-14 · Changhan Wang, Anne Wu, Juan Pino, Alexei Baevski 외

In this paper, we improve speech translation (ST) through effectively leveraging large quantities of unlabeled speech and text data in different and complementary ways. We explore both pretraining and self-training by us…

Language ModelingLanguage ModellingTranslation

Text-Free Image-to-Speech Synthesis Using Learned Segmental Units

2020-12-31 · ACL 2021 5 · Wei-Ning Hsu, David Harwath, Christopher Song, James Glass

In this paper we present the first model for directly synthesizing fluent, natural-sounding spoken audio captions for images that does not require natural language text as an intermediate representation or source of supe…

Image CaptioningSpeech SynthesisVisual Grounding

On the Role of Visual Cues in Audiovisual Speech Enhancement

2020-04-25 · Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald, Erik Marchi 외

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target spee…

Self-Supervised LearningSpeech Enhancement

LiRA: Learning Visual Speech Representations from Audio through Self-supervision

2021-06-16 · Pingchuan Ma, Rodrigo Mira, Stavros Petridis, Björn W. Schuller 외

The large amount of audiovisual content being shared online today has drawn substantial attention to the prospect of audiovisual self-supervised learning. Recent works have focused on each of these modalities separately,…

Lip ReadingSelf-Supervised LearningSentence