paper-with-me

홈 › Papers

FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition

2025-06-12 · Jongsuk Kim, Jaemyung Yu, Minchan Kwon, Junmo Kim

Large-scale ASR models have achieved remarkable gains in accuracy and robustness. However, fairness issues remain largely unaddressed despite their critical importance in real-world applications. In this work, we introduce FairASR, a system that mitigates demographic bias by learning representations that are uninformative about group membership, enabling fair generalization across demographic groups. Leveraging a multi-demographic dataset, our approach employs a gradient reversal layer to suppress demographic-discriminative features while maintaining the ability to capture generalizable speech patterns through an unsupervised contrastive loss. Experimental results show that FairASR delivers competitive overall ASR performance while significantly reducing performance disparities across different demographic groups.

📄 PDF Abstract BibTeX arXiv:2506.10747

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionContrastive LearningFairnessspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Few-Shot Contrastive Adaptation for Audio Abuse Detection in Low-Resource Indic Languages

2026-04-10 · Aditya Narayan Sankaran, Reza Farahbakhsh, Noel Crespi arxiv

Abusive speech detection is becoming increasingly important as social media shifts towards voice-based interaction, particularly in multilingual and low-resource settings. Most current systems rely on automatic speech re…

Speech Recognition

Learning Speech Representation From Contrastive Token-Acoustic Pretraining

2023-09-01 · Chunyu Qiang, Hao Li, Yixin Tian, Ruibo Fu 외

For fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate representations extracted from speech s…

Audio ClassificationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive Learning+6

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

2024-01-07 · Qiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu 외

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is su…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive Learning+6

Automatic Data Augmentation Selection and Parametrization in Contrastive Self-Supervised Speech Representation Learning

2022-04-08 · Salah Zaiem, Titouan Parcollet, Slim Essid

Contrastive learning enables learning useful audio and speech representations without ground-truth labels by maximizing the similarity between latent representations of similar signal segments. In this framework various …

Contrastive LearningData AugmentationRepresentation LearningSpeech Representation Learning

Prompting Audios Using Acoustic Properties For Emotion Representation

2023-10-03 · Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh, Huaming Wang 외

Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propos…

Contrastive LearningDiversityEmotion RecognitionRetrieval+1