Improving Fairness and Robustness in End-to-End Speech Recognition through unsupervised clustering
The challenge of fairness arises when Automatic Speech Recognition (ASR) systems do not perform equally well for all sub-groups of the population. In the past few years there have been many improvements in overall speech recognition quality, but without any particular focus on advancing Equality and Equity for all user groups for whom systems do not perform well. ASR fairness is therefore also a robustness issue. Meanwhile, data privacy also takes priority in production systems. In this paper, we present a privacy preserving approach to improve fairness and robustness of end-to-end ASR without using metadata, zip codes, or even speaker or utterance embeddings directly in training. We extract utterance level embeddings using a speaker ID model trained on a public dataset, which we then use in an unsupervised fashion to create acoustic clusters. We use cluster IDs instead of speaker utterance embeddings as extra features during model training, which shows improvements for all demographic groups and in particular for different accents.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringFairnessPrivacy Preservingspeech-recognitionSpeech RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
While subgroup disparities and performance bias are increasingly studied in computational research, fairness in categorical Speech Emotion Recognition (SER) remains underexplored. Existing methods often rely on explicit …
Emotion RecognitionFairnessSpeech Emotion RecognitionFairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
Large-scale ASR models have achieved remarkable gains in accuracy and robustness. However, fairness issues remain largely unaddressed despite their critical importance in real-world applications. In this work, we introdu…
Automatic Speech RecognitionContrastive LearningFairnessspeech-recognition+1Analyzing the Robustness of Unsupervised Speech Recognition
Unsupervised speech recognition (unsupervised ASR) aims to learn the ASR system with non-parallel speech and text corpus only. Wav2vec-U has shown promising results in unsupervised ASR by self-supervised speech represent…
Generative Adversarial Networkspeech-recognitionSpeech RecognitionUnsupervised Speech RecognitionTesting Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
Machine learning models for speech emotion recognition (SER) can be trained for different tasks and are usually evaluated based on a few available datasets per task. Tasks could include arousal, valence, dominance, emoti…
Emotion RecognitionFairnessSpeech Emotion RecognitionFairness of Automatic Speech Recognition in Cleft Lip and Palate Speech
Speech produced by individuals with cleft lip and palate (CLP) is often highly nasalized and breathy due to structural anomalies, causing shifts in formant structure that affect automatic speech recognition (ASR) perform…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Fairnessspeech-recognition+1