Who Decides if AI is Fair? The Labels Problem in Algorithmic Auditing
Labelled "ground truth" datasets are routinely used to evaluate and audit AI algorithms applied in high-stakes settings. However, there do not exist widely accepted benchmarks for the quality of labels in these datasets. We provide empirical evidence that quality of labels can significantly distort the results of algorithmic audits in real-world settings. Using data annotators typically hired by AI firms in India, we show that fidelity of the ground truth data can lead to spurious differences in performance of ASRs between urban and rural populations. After a rigorous, albeit expensive, label cleaning process, these disparities between groups disappear. Our findings highlight how trade-offs between label quality and data annotation costs can complicate algorithmic audits in practice. They also emphasize the need for development of consensus-driven, widely accepted benchmarks for label quality.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Peer-induced Fairness: A Causal Approach for Algorithmic Fairness Auditing
With the European Union's Artificial Intelligence Act taking effect on 1 August 2024, high-risk AI applications must adhere to stringent transparency and fairness standards. This paper addresses a crucial question: how c…
Causal InferencecounterfactualDecision MakingFairnessAssessing Classifier Fairness with Collider Bias
The increasing application of machine learning techniques in everyday decision-making processes has brought concerns about the fairness of algorithmic decision-making. This paper concerns the problem of collider bias whi…
Decision MakingFairnessAuditing LLMs for Algorithmic Fairness in Casenote-Augmented Tabular Prediction
LLMs are increasingly being considered for prediction tasks in high-stakes social service settings, but their algorithmic fairness properties in this context are poorly understood. In this short technical report, we audi…
Multi-class ClassificationMathematical Framework for Online Social Media Auditing
Social media platforms (SMPs) leverage algorithmic filtering (AF) as a means of selecting the content that constitutes a user's feed with the aim of maximizing their rewards. Selectively choosing the contents to be shown…
Decision MakingAuditing and Enforcing Conditional Fairness via Optimal Transport
Conditional demographic parity (CDP) is a measure of the demographic parity of a predictive model or decision process when conditioning on an additional feature or set of features. Many algorithmic fairness techniques ex…
Fairness