paper-with-me

홈 › Papers

Learning Noise-Invariant Representations for Robust Speech Recognition

2018-10-02 · Anonymous

Despite rapid advances in speech recognition, current models remain brittle to superficial perturbations to their inputs. Small amounts of noise can destroy the performance of an otherwise state-of-the-art model. To harden models against background noise, practitioners often perform data augmentation, adding artificially-noised examples to the training set, carrying over the original label. In this paper, we hypothesize that a clean example and its superficially perturbed counterparts shouldn't merely map to the same class--- they should map to the same representation. We propose invariant-representation-learning (IRL): At each training iteration, for each training example, we sample a noisy counterpart. We then apply a penalty term to coerce matched representations at each layer (above some chosen layer). Our key results, demonstrated on the LibriSpeech dataset are the following: (i) IRL significantly reduces character error rates (CER)on both clean' (3.3% vs 6.5%) and other' (11.0% vs 18.1%) test sets; (ii) on several out-of-domain noise settings (different from those seen during training), IRL's benefits are even more pronounced. Careful ablations confirm that our results are not simply due to shrinking activations at the chosen layers.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationRepresentation LearningRobust Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning

2022-12-10 · Chen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou 외

Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant re…

Audio-Visual Speech Recognitionreinforcement-learningReinforcement Learning (RL)speech-recognition+2

Invariant Representations for Noisy Speech Recognition

2016-11-27 · Dmitriy Serdyuk, Kartik Audhkhasi, Philémon Brakel, Bhuvana Ramabhadran 외

Modern automatic speech recognition (ASR) systems need to be robust under acoustic variability arising from environmental, speaker, channel, and recording conditions. Ensuring such robustness to variability is a challeng…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain AdaptationImage Generation+3

Learning Noise-Invariant Representations for Robust Speech Recognition

2018-07-17 · Davis Liang, Zhiheng Huang, Zachary C. Lipton

Despite rapid advances in speech recognition, current models remain brittle to superficial perturbations to their inputs. Small amounts of noise can destroy the performance of an otherwise state-of-the-art model. To hard…

Data AugmentationRepresentation LearningRobust Speech Recognitionspeech-recognition+1

Supervised Contrastive Learning for Accented Speech Recognition

2021-07-02 · Tao Han, Hantao Huang, Ziang Yang, Wei Han

Neural network based speech recognition systems suffer from performance degradation due to accented speech, especially unfamiliar accents. In this paper, we study the supervised contrastive learning framework for accente…

Accented Speech RecognitionContrastive LearningData AugmentationSentence+2

Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

2023-05-16 · Yuchen Hu, Ruizhe Li, Chen Chen, Heqing Zou 외

Audio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-invariant visual information. However, mos…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+2