paper-with-me

Papers

Multimodal Speech Recognition with Unstructured Audio Masking

2020-10-16 · EMNLP (nlpbt) 2020 11 · Tejas Srinivasan, Ramon Sanabria, Florian Metze, Desmond Elliott

Visual context has been shown to be useful for automatic speech recognition (ASR) systems when the speech signal is noisy or corrupted. Previous work, however, has only demonstrated the utility of visual context in an unrealistic setting, where a fixed set of words are systematically masked in the audio. In this paper, we simulate a more realistic masking scenario during model training, called RandWordMask, where the masking can occur for any word segment. Our experiments on the Flickr 8K Audio Captions Corpus show that multimodal ASR can generalize to recover different types of masked words in this unstructured masking setting. Moreover, our analysis shows that our models are capable of attending to the visual signal when the audio signal is corrupted. These results show that multimodal ASR systems can leverage the visual signal in more generalized noisy scenarios.

📄 PDF Abstract BibTeX arXiv:2010.08642

Code (0)

등록된 구현이 없습니다.

Tasks

8kAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Learning Contextually Fused Audio-visual Representations for Audio-visual Speech Recognition

2022-02-15 · Zi-Qiang Zhang, Jie Zhang, Jian-Shu Zhang, Ming-Hui Wu 외

With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech re…

Audio-Visual Speech RecognitionLipreadingRepresentation LearningSelf-Supervised Learning+3

Improving Multimodal Speech Recognition by Data Augmentation and Speech Representations

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated to the audio input. While previous approa…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+1

MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization

2024-06-25 · Adriana Fernandez-Lopez, Honglie Chen, Pingchuan Ma, Lu Yin 외

Pre-trained models have been a foundational approach in speech recognition, albeit with associated additional costs. In this study, we propose a regularization technique that facilitates the training of visual and audio-…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

SpliceOut: A Simple and Efficient Audio Augmentation Method

2021-09-30 · Arjit Jain, Pranay Reddy Samala, Deepak Mittal, Preethi Jyoti 외

Time masking has become a de facto augmentation technique for speech and audio tasks, including automatic speech recognition (ASR) and audio classification, most notably as a part of SpecAugment. In this work, we propose…

Audio ClassificationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Music Classification+4

MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

2025-03-14 · Jeong Hun Yeo, Hyeongseop Rha, Se Jin Park, Yong Man Ro

Audio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information. However, recent Large Language Model (LLM) based AVSR systems incur high compu…

Audio-Visual Speech RecognitionComputational EfficiencyLanguage ModelingLanguage Modelling+5