paper-with-me

홈 › Papers

Self-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling

2022-02-07 · Puyuan Peng, David Harwath

In this paper, we describe our submissions to the ZeroSpeech 2021 Challenge and SUPERB benchmark. Our submissions are based on the recently proposed FaST-VGS model, which is a Transformer-based model that learns to associate raw speech waveforms with semantically related images, all without the use of any transcriptions of the speech. Additionally, we introduce a novel extension of this model, FaST-VGS+, which is learned in a multi-task fashion with a masked language modeling objective in addition to the visual grounding objective. On ZeroSpeech 2021, we show that our models perform competitively on the ABX task, outperform all other concurrent submissions on the Syntactic and Semantic tasks, and nearly match the best system on the Lexical task. On the SUPERB benchmark, we show that our models also achieve strong performance, in some cases even outperforming the popular wav2vec2.0 model.

📄 PDF Abstract BibTeX arXiv:2202.03543

Code (1)

jasonppy/fast-vgs-family 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingMasked Language ModelingRepresentation LearningVisual Grounding

Similar Papers 제목 키워드 기반

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

2025-09-22 · María Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi, Maureen de Seyssel arxiv

Self-supervised learning (SSL) has made significant advances in speech representation learning. Models like wav2vec 2.0 and HuBERT have achieved state-of-the-art results in tasks such as speech recognition, particularly …

Self-Supervised LearningRepresentation LearningSpeech RecognitionVisual Grounding

Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding

2026-08-27 · Gabriel Pirlogeanu, Dan Oneata, Horia Cucu, Herman Kamper arxiv

In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speec…

Visual GroundingImage CaptioningKeyword Spotting

Syllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model

2023-05-19 · Puyuan Peng, Shang-Wen Li, Okko Räsänen, Abdelrahman Mohamed 외

In this paper, we show that representations capturing syllabic units emerge when training a self-supervised speech model with a visually-grounded training objective. We demonstrate that a nearly identical model architect…

Language ModelingLanguage ModellingMasked Language ModelingSegmentation+2

Text-Free Image-to-Speech Synthesis Using Learned Segmental Units

2020-12-31 · ACL 2021 5 · Wei-Ning Hsu, David Harwath, Christopher Song, James Glass

In this paper we present the first model for directly synthesizing fluent, natural-sounding spoken audio captions for images that does not require natural language text as an intermediate representation or source of supe…

Image CaptioningSpeech SynthesisVisual Grounding

Bridging the Gap: Using Deep Acoustic Representations to Learn Grounded Language from Percepts and Raw Speech

2021-12-27 · Gaoussou Youssouf Kebe, Luke E. Richards, Edward Raff, Francis Ferraro 외

Learning to understand grounded language, which connects natural language to percepts, is a critical research area. Prior work in grounded language acquisition has focused primarily on textual inputs. In this work we dem…

Language Acquisitionspeech-recognitionSpeech Recognition