Evaluating and Improving the Robustness of Speech Command Recognition Models to Noise and Distribution Shifts
Although prior work in computer vision has shown strong correlations between in-distribution (ID) and out-of-distribution (OOD) accuracies, such relationships remain underexplored in audio-based models. In this study, we investigate how training conditions and input features affect the robustness and generalization abilities of spoken keyword classifiers under OOD conditions. We benchmark several neural architectures across a variety of evaluation sets. To quantify the impact of noise on generalization, we make use of two metrics: Fairness (F), which measures overall accuracy gains compared to a baseline model, and Robustness (R), which assesses the convergence between ID and OOD performance. Our results suggest that noise-aware training improves robustness in some configurations. These findings shed new light on the benefits and limitations of noise-based augmentation for generalization in speech models.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition
We present an MatchboxNet - an end-to-end neural network for speech command recognition. MatchboxNet is a deep residual network composed from blocks of 1D time-channel separable convolution, batch-normalization, ReLU and…
Data AugmentationTime Series AnalysisTowards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer
In light of the widespread application of Automatic Speech Recognition (ASR) systems, their security concerns have received much more attention than ever before, primarily due to the susceptibility of Deep Neural Network…
Adversarial AttackAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+4Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study
Noise robustness is a key aspect of successful speech applications. Speech enhancement (SE) has been investigated to improve automatic speech recognition accuracy; however, its effectiveness for keyword spotting (KWS) is…
Automatic Speech RecognitionKeyword SpottingSpeech Enhancementspeech-recognition+1Sisyphus, a Workflow Manager Designed for Machine Translation and Automatic Speech Recognition
Training and testing many possible parameters or model architectures of state-of-the-art machine translation or automatic speech recognition system is a cumbersome task. They usually require a long pipeline of commands r…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+2Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an e…
Acoustic Scene ClassificationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Recognition+4