AequeVox: Automated Fairness Testing of Speech Recognition Systems
Automatic Speech Recognition (ASR) systems have become ubiquitous. They can be found in a variety of form factors and are increasingly important in our daily lives. As such, ensuring that these systems are equitable to different subgroups of the population is crucial. In this paper, we introduce, AequeVox, an automated testing framework for evaluating the fairness of ASR systems. AequeVox simulates different environments to assess the effectiveness of ASR systems for different populations. In addition, we investigate whether the chosen simulations are comprehensible to humans. We further propose a fault localization technique capable of identifying words that are not robust to these varying environments. Both components of AequeVox are able to operate in the absence of ground truth data. We evaluated AequeVox on speech from four different datasets using three different commercial ASRs. Our experiments reveal that non-native English, female and Nigerian English speakers generate 109%, 528.5% and 156.9% more errors, on average than native English, male and UK Midlands speakers, respectively. Our user study also reveals that 82.9% of the simulations (employed through speech transformations) had a comprehensibility rating above seven (out of ten), with the lowest rating being 6.78. This further validates the fairness violations discovered by AequeVox. Finally, we show that the non-robust words, as predicted by the fault localization technique embodied in AequeVox, show 223.8% more errors than the predicted robust words across all ASRs.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)FairnessFault localizationspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Fairness of Automatic Speech Recognition in Cleft Lip and Palate Speech
Speech produced by individuals with cleft lip and palate (CLP) is often highly nasalized and breathy due to structural anomalies, causing shifts in formant structure that affect automatic speech recognition (ASR) perform…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Fairnessspeech-recognition+1Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
Machine learning models for speech emotion recognition (SER) can be trained for different tasks and are usually evaluated based on a few available datasets per task. Tasks could include arousal, valence, dominance, emoti…
Emotion RecognitionFairnessSpeech Emotion RecognitionToward Fairness in AI for People with Disabilities: A Research Roadmap
AI technologies have the potential to dramatically impact the lives of people with disabilities (PWD). Indeed, improving the lives of PWD is a motivator for many state-of-the-art AI systems, such as automated speech reco…
Fairnessspeech-recognitionSpeech RecognitionAutomated Testing of AI Models
The last decade has seen tremendous progress in AI technology and applications. With such widespread adoption, ensuring the reliability of the AI models is crucial. In past, we took the first step of creating a testing f…
FairnessSpeech-to-Texttext-classificationText Classification+2ASDF: A Differential Testing Framework for Automatic Speech Recognition Systems
Recent years have witnessed wider adoption of Automated Speech Recognition (ASR) techniques in various domains. Consequently, evaluating and enhancing the quality of ASR systems is of great importance. This paper propose…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition