paper-with-me

홈 › Papers

Image Counterfactual Sensitivity Analysis for Detecting Unintended Bias

2019-06-14 · Emily Denton, Ben Hutchinson, Margaret Mitchell, Timnit Gebru, Andrew Zaldivar

Facial analysis models are increasingly used in applications that have serious impacts on people's lives, ranging from authentication to surveillance tracking. It is therefore critical to develop techniques that can reveal unintended biases in facial classifiers to help guide the ethical use of facial analysis technology. This work proposes a framework called \textit{image counterfactual sensitivity analysis}, which we explore as a proof-of-concept in analyzing a smiling attribute classifier trained on faces of celebrities. The framework utilizes counterfactuals to examine how a classifier's prediction changes if a face characteristic slightly changes. We leverage recent advances in generative adversarial networks to build a realistic generative model of face images that affords controlled manipulation of specific image characteristics. We then introduce a set of metrics that measure the effect of manipulating a specific property on the output of the trained classifier. Empirically, we find several different factors of variation that affect the predictions of the smiling classifier. This proof-of-concept demonstrates potential ways generative models can be leveraged for fine-grained analysis of bias and fairness.

📄 PDF Abstract BibTeX arXiv:1906.06439

Code (0)

등록된 구현이 없습니다.

Tasks

AttributecounterfactualFairnessSensitivity

Methods 이 논문이 사용한 방법론

Counterfactuals 설명 없음

Similar Papers 제목 키워드 기반

Prediction Sensitivity: Continual Audit of Counterfactual Fairness in Deployed Classifiers

2022-02-09 · Krystal Maughan, Ivoline C. Ngong, Joseph P. Near

As AI-based systems increasingly impact many areas of our lives, auditing these systems for fairness is an increasingly high-stakes problem. Traditional group fairness metrics can miss discrimination against individuals …

counterfactualFairnessPredictionSensitivity

An Interpretable Local Editing Model for Counterfactual Medical Image Generation

2026-02-28 · Hyungi Min, Taeseung You, Hangyeul Lee, Yeongjae Cho 외 arxiv

Counterfactual medical image generation have emerged as a critical tool for enhancing AI-driven systems in medical domain by answering "what-if" questions. However, existing approaches face two fundamental limitations: F…

Medical Image Generation

Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection

2022-05-09 · NAACL 2022 7 · Indira Sen, Mattia Samory, Claudia Wagner, Isabelle Augenstein

Counterfactually Augmented Data (CAD) aims to improve out-of-domain generalizability, an indicator of model robustness. The improvement is credited with promoting core features of the construct over spurious artifacts th…

Hate Speech Detection

Sensitivity Analysis in Unconditional Quantile Effects

2023-03-24 · Julian Martinez-Iriarte

This paper proposes a framework to analyze the effects of counterfactual policies on the unconditional quantiles of an outcome variable. For a given counterfactual policy, we obtain identified sets for the effect of both…

counterfactualSelection biasSensitivity

Counterfactual Visual Explanation via Causally-Guided Adversarial Steering

2025-07-14 · Yiran Qiao, Disheng Liu, Yiren Lu, Yu Yin 외 arxiv

Recent work on counterfactual visual explanations has contributed to making artificial intelligence models more explainable by providing visual perturbation to flip the prediction. However, these approaches neglect the c…

Image Generation