Learning Noise-Invariant Representations for Robust Speech Recognition
Despite rapid advances in speech recognition, current models remain brittle to superficial perturbations to their inputs. Small amounts of noise can destroy the performance of an otherwise state-of-the-art model. To harden models against background noise, practitioners often perform data augmentation, adding artificially-noised examples to the training set, carrying over the original label. In this paper, we hypothesize that a clean example and its superficially perturbed counterparts shouldn't merely map to the same class --- they should map to the same representation. We propose invariant-representation-learning (IRL): At each training iteration, for each training example,we sample a noisy counterpart. We then apply a penalty term to coerce matched representations at each layer (above some chosen layer). Our key results, demonstrated on the Librispeech dataset are the following: (i) IRL significantly reduces character error rates (CER) on both 'clean' (3.3% vs 6.5%) and 'other' (11.0% vs 18.1%) test sets; (ii) on several out-of-domain noise settings (different from those seen during training), IRL's benefits are even more pronounced. Careful ablations confirm that our results are not simply due to shrinking activations at the chosen layers.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationRepresentation LearningRobust Speech Recognitionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning
Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant re…
Audio-Visual Speech Recognitionreinforcement-learningReinforcement Learning (RL)speech-recognition+2Invariant Representations for Noisy Speech Recognition
Modern automatic speech recognition (ASR) systems need to be robust under acoustic variability arising from environmental, speaker, channel, and recording conditions. Ensuring such robustness to variability is a challeng…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain AdaptationImage Generation+3Learning Noise-Invariant Representations for Robust Speech Recognition
Despite rapid advances in speech recognition, current models remain brittle to superficial perturbations to their inputs. Small amounts of noise can destroy the performance of an otherwise state-of-the-art model. To hard…
Data AugmentationRepresentation LearningRobust Speech Recognitionspeech-recognition+1Supervised Contrastive Learning for Accented Speech Recognition
Neural network based speech recognition systems suffer from performance degradation due to accented speech, especially unfamiliar accents. In this paper, we study the supervised contrastive learning framework for accente…
Accented Speech RecognitionContrastive LearningData AugmentationSentence+2Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition
Audio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-invariant visual information. However, mos…
Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+2