paper-with-me

홈 › Papers

Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset Transcriptions

2021-11-18 · Chunxi Liu, Michael Picheny, Leda Sari, Pooja Chitkara, Alex Xiao, Xiaohui Zhang, Mark Chou, Andres Alvarado, Caner Hazirbas, Yatharth Saraf

It is well known that many machine learning systems demonstrate bias towards specific groups of individuals. This problem has been studied extensively in the Facial Recognition area, but much less so in Automatic Speech Recognition (ASR). This paper presents initial Speech Recognition results on "Casual Conversations" -- a publicly released 846 hour corpus designed to help researchers evaluate their computer vision and audio models for accuracy across a diverse set of metadata, including age, gender, and skin tone. The entire corpus has been manually transcribed, allowing for detailed ASR evaluations across these metadata. Multiple ASR models are evaluated, including models trained on LibriSpeech, 14,000 hour transcribed, and over 2 million hour untranscribed social media videos. Significant differences in word error rate across gender and skin tone are observed at times for all models. We are releasing human transcripts from the Casual Conversations dataset to encourage the community to develop a variety of techniques to reduce these statistical biases.

📄 PDF Abstract BibTeX arXiv:2111.09983

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Fairnessspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

The Nijmegen Corpus of Casual Czech

2014-05-01 · LREC 2014 5 · Mirjam Ernestus, Lucie Ko{\v{c}}kov{\'a}-Amortov{\'a}, Petr Pollak

This article introduces a new speech corpus, the Nijmegen Corpus of Casual Czech (NCCCz), which contains more than 30 hours of high-quality recordings of casual conversations in Common Czech, among ten groups of three ma…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over

2024-06-27 · Atsunori Ogawa, Naoyuki Kamo, Kohei Matsuura, Takanori Ashihara 외

Large language models (LLMs) have been successfully applied for rescoring automatic speech recognition (ASR) hypotheses. However, their ability to rescore ASR hypotheses of casual conversations has not been sufficiently …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain Adaptationspeech-recognition+1

Casual Conversations v2: Designing a large consent-driven dataset to measure algorithmic bias and robustness

2022-11-10 · Caner Hazirbas, Yejin Bang, Tiezheng Yu, Parisa Assar 외

Developing robust and fair AI systems require datasets with comprehensive set of labels that can help ensure the validity and legitimacy of relevant measurements. Recent efforts, therefore, focus on collecting person-rel…

Fairness

Affect Recognition in Conversations Using Large Language Models

2023-09-22 · Shutong Feng, Guangzhi Sun, Nurul Lubis, Wen Wu 외

Affect recognition, encompassing emotions, moods, and feelings, plays a pivotal role in human communication. In the realm of conversational artificial intelligence, the ability to discern and respond to human affective c…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)In-Context Learningspeech-recognition+1

A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results

2025-10-27 · Thai-Binh Nguyen, Katerina Zmolikova, Pingchuan Ma, Ngoc Quan Pham 외 arxiv

We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a single-room setting using audio, visual, …