Reliability Gaps Between Groups in COMPAS Dataset
This paper investigates the inter-rater reliability of risk assessment instruments (RAIs). The main question is whether different, socially salient groups are affected differently by a lack of inter-rater reliability of RAIs, that is, whether mistakes with respect to different groups affects them differently. The question is investigated with a simulation study of the COMPAS dataset. A controlled degree of noise is injected into the input data of a predictive model; the noise can be interpreted as a synthetic rater that makes mistakes. The main finding is that there are systematic differences in output reliability between groups in the COMPAS dataset. The sign of the difference depends on the kind of inter-rater statistic that is used (Cohen's Kappa, Byrt's PABAK, ICC), and in particular whether or not a correction of predictions prevalences of the groups is used.
Code (1)
Similar Papers 제목 키워드 기반
Extended Stochastic Block Models with Application to Criminal Networks
Reliably learning group structures among nodes in network data is challenging in several applications. We are particularly motivated by studying covert networks that encode relationships among criminals. These data are s…
Community DetectionModel SelectionUncertainty QuantificationMultilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems
Conversational explainable artificial intelligence (ConvXAI) systems based on large language models (LLMs) have garnered considerable attention for their ability to enhance user comprehension through dialogue-based expla…
Intent RecognitionAssessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for Support
Various tools and practices have been developed to support practitioners in identifying, assessing, and mitigating fairness-related harms caused by AI systems. However, prior research has highlighted gaps between the int…
FairnessMeasuring Changes in Disparity Gaps: An Application to Health Insurance
We propose a method for reporting how program evaluations reduce gaps between groups, such as the gender or Black-white gap. We first show that the reduction in disparities between groups can be written as the difference…
Fair Dataset Distillation via Cross-Group Barycenter Alignment
Dataset Distillation aims to compress a large dataset into a small synthetic one while maintaining predictive performance. We show that as different demographic groups exhibit distinct predictive patterns, the distillati…