paper-with-me

홈 › Papers

Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models

2024-03-31 · Alkis Koudounas, Flavio Giobergia

The Fearless Steps APOLLO Community Resource provides unparalleled opportunities to explore the potential of multi-speaker team communications from NASA Apollo missions. This study focuses on discovering the characteristics that make Apollo recordings more or less intelligible to Automatic Speech Recognition (ASR) methods. We extract, for each audio recording, interpretable metadata on recordings (signal-to-noise ratio, spectral flatness, presence of pauses, sentence duration), transcript (number of words spoken, speaking rate), or known a priori (speaker). We identify subgroups of audio recordings based on combinations of these metadata and compute each subgroup's performance (e.g., Word Error Rate) and the difference in performance (''divergence'') w.r.t the overall population. We then apply the Whisper model in different sizes, trained on English-only or multilingual datasets, in zero-shot or after fine-tuning. We conduct several analyses to (i) automatically identify and describe the most problematic subgroups for a given model, (ii) examine the impact of fine-tuning w.r.t. zero-shot at the subgroup level, (iii) understand the effect of model size on subgroup performance, and (iv) analyze if multilingual models are more sensitive than monolingual to subgroup performance disparities. The insights enhance our understanding of subgroup-specific performance variations, paving the way for advancements in optimizing ASR systems for Earth-to-space communications.

📄 PDF Abstract BibTeX arXiv:2404.07226

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Apollo Please enter a description about the method here

Similar Papers 제목 키워드 기반

Prompt Fairness: Sub-group Disparities in LLMs

2025-11-25 · Meiyu Zhong, Noel Teku, Ravi Tandon arxiv

Large Language Models (LLMs), though shown to be effective in many applications, can vary significantly in their response quality. In this paper, we investigate this problem of prompt fairness: specifically, the phrasing…

Identifying Biased Subgroups in Ranking and Classification

2021-08-17 · Eliana Pastor, Luca de Alfaro, Elena Baralis

When analyzing the behavior of machine learning algorithms, it is important to identify specific data subgroups for which the considered algorithm shows different performance with respect to the entire dataset. The inter…

Classification

Learning Exceptional Subgroups by End-to-End Maximizing KL-divergence

2024-02-20 · Sascha Xu, Nils Philipp Walter, Janis Kalofolias, Jilles Vreeken

Finding and describing sub-populations that are exceptional regarding a target property has important applications in many scientific disciplines, from identifying disadvantaged demographic groups in census data to findi…

Subgroup Performance Analysis in Hidden Stratifications

2025-03-13 · Alceu Bissoto, Trung-Dung Hoang, Tim Flühmann, Susu Sun 외

Machine learning (ML) models may suffer from significant performance disparities between patient groups. Identifying such disparities by monitoring performance at a granular level is crucial for safely deploying ML to ea…

Lesion ClassificationSkin Lesion ClassificationSubgroup Discovery

The Forest Behind the Tree: Heterogeneity in How US Governor's Party Affects Black Workers

2021-10-01 · Guy Tchuente, Johnson Kakeu, John Nana Francois

Income inequality is a distributional phenomenon. This paper examines the impact of U.S governor's party allegiance (Republican vs Democrat) on ethnic wage gap. A descriptive analysis of the distribution of yearly earnin…

Descriptive