paper-with-me

Papers

Visually Guided Self Supervised Learning of Speech Representations

2020-01-13 · Abhinav Shukla, Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis, Maja Pantic

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very limited work that studies the interaction between the two modalities for learning self supervised representations. We propose a framework for learning audio representations guided by the visual modality in the context of audiovisual speech. We employ a generative audio-to-video training scheme in which we animate a still image corresponding to a given audio clip and optimize the generated video to be as close as possible to the real video of the speech segment. Through this process, the audio encoder network learns useful speech representations that we evaluate on emotion recognition and speech recognition. We achieve state of the art results for emotion recognition and competitive results for speech recognition. This demonstrates the potential of visual supervision for learning audio representations as a novel way for self-supervised learning which has not been explored in the past. The proposed unsupervised audio features can leverage a virtually unlimited amount of training data of unlabelled audiovisual speech and have a large number of potentially promising applications.

📄 PDF Abstract BibTeX arXiv:2001.04316

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionRepresentation LearningSelf-Supervised LearningSpeech Emotion Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Syllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model

2023-05-19 · Puyuan Peng, Shang-Wen Li, Okko Räsänen, Abdelrahman Mohamed 외

In this paper, we show that representations capturing syllabic units emerge when training a self-supervised speech model with a visually-grounded training objective. We demonstrate that a nearly identical model architect…

Language ModelingLanguage ModellingMasked Language ModelingSegmentation+2

Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks

2024-02-24 · Duo Ma, Xianghu Yue, Junyi Ao, Xiaoxue Gao 외

Human language can be expressed in either written or spoken form, i.e. text or speech. Humans can acquire knowledge from text to improve speaking and listening. However, the quest for speech pre-trained models to leverag…

Pseudo LabelSelf-Supervised Learning

Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System

2023-06-05 · Khazar Khorrami, María Andrea Cruz Blandón, Tuomas Virtanen, Okko Räsänen

Speech representation learning with self-supervised algorithms has resulted in notable performance boosts in many downstream tasks. Recent work combined self-supervised learning (SSL) and visually grounded speech (VGS) p…

Multi-Task LearningRepresentation LearningRetrievalSelf-Supervised Learning+2

Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations

2026-01-28 · Amit Meghanani, Thomas Hain arxiv

Integrating front-end speech enhancement (SE) models with self-supervised learning (SSL)-based speech models is effective for downstream tasks in noisy conditions. SE models are commonly fine-tuned using SSL representati…

Self-Supervised LearningSpeech Enhancement

Integrating Self-supervised Speech Model with Pseudo Word-level Targets from Visually-grounded Speech Model

2024-02-08 · Hung-Chieh Fang, Nai-Xuan Ye, Yi-Jen Shih, Puyuan Peng 외

Recent advances in self-supervised speech models have shown significant improvement in many downstream tasks. However, these models predominantly centered on frame-level training objectives, which can fall short in spoke…

modelSpoken Language Understanding