paper-with-me

홈 › Papers

Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emotion Recognition

2023-10-29 · Isaac Slaughter, Craig Greenberg, Reva Schwartz, Aylin Caliskan

Previous work has established that a person's demographics and speech style affect how well speech processing models perform for them. But where does this bias come from? In this work, we present the Speech Embedding Association Test (SpEAT), a method for detecting bias in one type of model used for many speech tasks: pre-trained models. The SpEAT is inspired by word embedding association tests in natural language processing, which quantify intrinsic bias in a model's representations of different concepts, such as race or valence (something's pleasantness or unpleasantness) and capture the extent to which a model trained on large-scale socio-cultural data has learned human-like biases. Using the SpEAT, we test for six types of bias in 16 English speech models (including 4 models also trained on multilingual data), which come from the wav2vec 2.0, HuBERT, WavLM, and Whisper model families. We find that 14 or more models reveal positive valence (pleasantness) associations with abled people over disabled people, with European-Americans over African-Americans, with females over males, with U.S. accented speakers over non-U.S. accented speakers, and with younger people over older people. Beyond establishing that pre-trained speech models contain these biases, we also show that they can have real world effects. We compare biases found in pre-trained models to biases in downstream models adapted to the task of Speech Emotion Recognition (SER) and find that in 66 of the 96 tests performed (69%), the group that is more associated with positive valence as indicated by the SpEAT also tends to be predicted as speaking with higher valence by the downstream model. Our work provides evidence that, like text and image-based models, pre-trained speech based-models frequently learn human-like biases. Our work also shows that bias found in pre-trained models can propagate to the downstream task of SER.

📄 PDF Abstract BibTeX arXiv:2310.18877

Code (1)

isaaconline/speat 공식 구현

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

Speech Denoising Convolutional Neural Network trained with Deep Feature Losses.

2018-06-27 · Interspeech 2018 6 · Francois G. Germain, Qifeng Chen, Vladlen Koltun

We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to prod…

Audio TaggingDenoisingSpeech DenoisingSpeech Enhancement

Speech Denoising with Deep Feature Losses

2018-06-27 · Francois G. Germain, Qifeng Chen, Vladlen Koltun

We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly. Given input audio containing speech corrupted by an additive background signal, the system aims to prod…

Audio TaggingDenoisingSpeech Denoising

A Study of Gender Impact in Self-supervised Models for Speech-to-Text Systems

2022-04-04 · Marcely Zanon Boito, Laurent Besacier, Natalia Tomashenko, Yannick Estève

Self-supervised models for speech processing emerged recently as popular foundation blocks in speech processing pipelines. These models are pre-trained on unlabeled audio data and then used in speech processing downstrea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Fairnessspeech-recognition+2

CAFE A Novel Code switching Dataset for Algerian Dialect French and English

2024-11-20 · Houssam Eddine-Othman Lachemat, Akli Abbas, Nourredine Oukas, Yassine El Kheir 외

The paper introduces and publicly releases (Data download link available after acceptance) CAFE -- the first Code-switching dataset between Algerian dialect, French, and english languages. The CAFE speech data is unique …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Pseudo Labelspeech-recognition+1

BASPRO: a balanced script producer for speech corpus collection based on the genetic algorithm

2022-12-11 · Yu-Wen Chen, Hsin-Min Wang, Yu Tsao

The performance of speech-processing models is heavily influenced by the speech corpus that is used for training and evaluation. In this study, we propose BAlanced Script PROducer (BASPRO) system, which can automatically…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)SentenceSpeech Enhancement+4