paper-with-me

홈 › Papers

When Fairness Metrics Disagree: Evaluating the Reliability of Demographic Fairness Assessment in Machine Learning

2026-04-16 · Khalid Adnan Alsayed arxiv

The evaluation of fairness in machine learning systems has become a central concern in high-stakes applications, including biometric recognition, healthcare decision-making, and automated risk assessment. Existing approaches typically rely on a small number of fairness metrics to assess model behaviour across group partitions, implicitly assuming that these metrics provide consistent and reliable conclusions. However, different fairness metrics capture distinct statistical properties of model performance and may therefore produce conflicting assessments when applied to the same system. In this work, we investigate the consistency of fairness evaluation by conducting a systematic multi-metric analysis of demographic bias in machine learning models. Using face recognition as a controlled experimental setting, we evaluate model performance across multiple group partitions under a range of commonly used fairness metrics, including error-rate disparities and performance-based measures. Our results demonstrate that fairness assessments can vary significantly depending on the choice of metrics, leading to contradictory conclusions regarding model bias. To quantify this phenomenon, we introduce the Fairness Disagreement Index (FDI), a measure designed to capture the degree of inconsistency across fairness metrics. We further show that disagreement remains high across thresholds and model configurations. These findings highlight a critical limitation in current fairness evaluation practices and suggest that single-metric reporting is insufficient for reliable bias assessment.

📄 PDF Abstract BibTeX arXiv:2604.15038

Code (0)

등록된 구현이 없습니다.

Tasks

Face Recognition

Similar Papers 제목 키워드 기반

Promises and Challenges of Causality for Ethical Machine Learning

2022-01-26 · Aida Rahmattalabi, Alice Xiang

In recent years, there has been increasing interest in causal reasoning for designing fair decision-making systems due to its compatibility with legal frameworks, interpretability for human stakeholders, and robustness t…

BIG-bench Machine LearningCausal InferenceDecision MakingEconometrics+1

Bursting the Burden Bubble? An Assessment of Sharma et al.'s Counterfactual-based Fairness Metric

2022-11-21 · Yochem van Rosmalen, Florian van der Steen, Sebastiaan Jans, Daan van der Weijden

Machine learning has seen an increase in negative publicity in recent years, due to biased, unfair, and uninterpretable models. There is a rising interest in making machine learning models more fair for unprivileged comm…

AttributecounterfactualFairness

Why Don't Prompt-Based Fairness Metrics Correlate?

2024-06-09 · Abdelrahman Zayed, Goncalo Mordido, Ioana Baldini, Sarath Chandar

The widespread use of large language models has brought up essential questions about the potential biases these models might learn. This led to the development of several metrics aimed at evaluating and mitigating these …

Fairness

Metrics also Disagree in the Low Scoring Range: Revisiting Summarization Evaluation Metrics

2020-11-08 · COLING 2020 8 · Manik Bhandari, Pranav Gour, Atabak Ashfaq, PengFei Liu

In text summarization, evaluating the efficacy of automatic metrics without human judgments has become recently popular. One exemplar work concludes that automatic metrics strongly disagree when ranking high-scoring summ…

Text Summarization

EXAGREE: Towards Explanation Agreement in Explainable Machine Learning

2024-11-04 · Sichao Li, Quanling Deng, Amanda S. Barnard

Explanations in machine learning are critical for trust, transparency, and fairness. Yet, complex disagreements among these explanations limit the reliability and applicability of machine learning models, especially in h…

Fairness