Transfer Learning from Audio-Visual Grounding to Speech Recognition
Transfer learning aims to reduce the amount of data required to excel at a new task by re-using the knowledge acquired from learning other related tasks. This paper proposes a novel transfer learning scenario, which distills robust phonetic features from grounding models that are trained to tell whether a pair of image and speech are semantically correlated, without using any textual transcripts. As semantics of speech are largely determined by its lexical content, grounding models learn to preserve phonetic information while disregarding uncorrelated factors, such as speaker and channel. To study the properties of features distilled from different layers, we use them as input separately to train multiple speech recognition models. Empirical results demonstrate that layers closer to input retain more phonetic information, while following layers exhibit greater invariance to domain shift. Moreover, while most previous studies include training data for speech recognition for feature extractor training, our grounding models are not trained on any of those data, indicating more universal applicability to new domains.
Code (0)
등록된 구현이 없습니다.
Tasks
speech-recognitionSpeech RecognitionTransfer LearningVisual GroundingSimilar Papers 제목 키워드 기반
You Only Speak Once to See
Grounding objects in images using visual cues is a well-established approach in computer vision, yet the potential of audio as a modality for object recognition and grounding remains underexplored. We introduce YOSS, "Yo…
Contrastive LearningObjectObject RecognitionScene UnderstandingAudio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D…
Multi-Label ClassificationRepresentation LearningSpeech RecognitionVisual GroundingFine-Grained Grounding for Multimodal Speech Recognition
Multimodal automatic speech recognition systems integrate information from images to improve speech recognition quality, by grounding the speech in the visual context. While visual signals have been shown to be useful fo…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionLeveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
Self-supervised learning (SSL) has made significant advances in speech representation learning. Models like wav2vec 2.0 and HuBERT have achieved state-of-the-art results in tasks such as speech recognition, particularly …
Self-Supervised LearningRepresentation LearningSpeech RecognitionVisual GroundingHearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition
Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most existing efforts still focus on audio modali…
Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition