paper-with-me

홈 › Papers

Transfer Learning from Audio-Visual Grounding to Speech Recognition

2019-07-09 · Wei-Ning Hsu, David Harwath, James Glass

Transfer learning aims to reduce the amount of data required to excel at a new task by re-using the knowledge acquired from learning other related tasks. This paper proposes a novel transfer learning scenario, which distills robust phonetic features from grounding models that are trained to tell whether a pair of image and speech are semantically correlated, without using any textual transcripts. As semantics of speech are largely determined by its lexical content, grounding models learn to preserve phonetic information while disregarding uncorrelated factors, such as speaker and channel. To study the properties of features distilled from different layers, we use them as input separately to train multiple speech recognition models. Empirical results demonstrate that layers closer to input retain more phonetic information, while following layers exhibit greater invariance to domain shift. Moreover, while most previous studies include training data for speech recognition for feature extractor training, our grounding models are not trained on any of those data, indicating more universal applicability to new domains.

📄 PDF Abstract BibTeX arXiv:1907.04355

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionTransfer LearningVisual Grounding

Similar Papers 제목 키워드 기반

You Only Speak Once to See

2024-09-27 · Wenhao Yang, Jianguo Wei, Wenhuan Lu, Lei LI

Grounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remains underexplored. We introduce YOSS, "Yo…

Contrastive LearningObjectObject RecognitionScene Understanding

Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding

2025-07-01 · Duc Cao-Dinh, Khai Le-Duc, Anh Dao, Bach Phan Tat 외 arxiv

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D…

Multi-Label ClassificationRepresentation LearningSpeech RecognitionVisual Grounding

Fine-Grained Grounding for Multimodal Speech Recognition

2020-10-05 · Findings of the Association for Computational Linguistics 2020 · Tejas Srinivasan, Ramon Sanabria, Florian Metze, Desmond Elliott

Multimodal automatic speech recognition systems integrate information from images to improve speech recognition quality, by grounding the speech in the visual context. While visual signals have been shown to be useful fo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

2025-09-22 · María Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi, Maureen de Seyssel arxiv

Self-supervised learning (SSL) has made significant advances in speech representation learning. Models like wav2vec 2.0 and HuBERT have achieved state-of-the-art results in tasks such as speech recognition, particularly …

Self-Supervised LearningRepresentation LearningSpeech RecognitionVisual Grounding

Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition

2023-06-18 · Yuchen Hu, Ruizhe Li, Chen Chen, Chengwei Qin 외

Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most existing efforts still focus on audio modali…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition