paper-with-me

Papers

Fairness via Representation Neutralization

2021-06-23 · NeurIPS 2021 12 · Mengnan Du, Subhabrata Mukherjee, Guanchu Wang, Ruixiang Tang, Ahmed Hassan Awadallah, Xia Hu

Existing bias mitigation methods for DNN models primarily work on learning debiased encoders. This process not only requires a lot of instance-level annotations for sensitive attributes, it also does not guarantee that all fairness sensitive information has been removed from the encoder. To address these limitations, we explore the following research question: Can we reduce the discrimination of DNN models by only debiasing the classification head, even with biased representations as inputs? To this end, we propose a new mitigation technique, namely, Representation Neutralization for Fairness (RNF) that achieves fairness by debiasing only the task-specific classification head of DNN models. To this end, we leverage samples with the same ground-truth label but different sensitive attributes, and use their neutralized representations to train the classification head of the DNN model. The key idea of RNF is to discourage the classification head from capturing spurious correlation between fairness sensitive information in encoder representations with specific class labels. To address low-resource settings with no access to sensitive attribute annotations, we leverage a bias-amplified model to generate proxy annotations for sensitive attributes. Experimental results over several benchmark datasets demonstrate our RNF framework to effectively reduce discrimination of DNN models with minimal degradation in task-specific performance.

📄 PDF Abstract BibTeX arXiv:2106.12674

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeClassificationFairness

Similar Papers 제목 키워드 기반

FairSIN: Achieving Fairness in Graph Neural Networks through Sensitive Information Neutralization

2024-03-19 · Cheng Yang, Jixi Liu, Yunhe Yan, Chuan Shi

Despite the remarkable success of graph neural networks (GNNs) in modeling graph-structured data, like other machine learning models, GNNs are also susceptible to making biased predictions based on sensitive attributes, …

Fairness

FairCLIP: Social Bias Elimination based on Attribute Prototype Learning and Representation Neutralization

2022-10-26 · Junyang Wang, Yi Zhang, Jitao Sang

The Vision-Language Pre-training (VLP) models like CLIP have gained popularity in recent years. However, many works found that the social biases hidden in CLIP easily manifest in downstream tasks, especially in image ret…

AttributeFairnessImage RetrievalRetrieval

Bias Neutralization Framework: Measuring Fairness in Large Language Models with Bias Intelligence Quotient (BiQ)

2024-04-28 · Malur Narayan, John Pasmore, Elton Sampaio, Vijay Raghavan 외

The burgeoning influence of Large Language Models (LLMs) in shaping public discourse and decision-making underscores the imperative to address inherent biases within these AI systems. In the wake of AI's expansive integr…

Decision MakingFairnessLanguage ModelingLanguage Modelling+1

Prompt Fairness: Sub-group Disparities in LLMs

2025-11-25 · Meiyu Zhong, Noel Teku, Ravi Tandon arxiv

Large Language Models (LLMs), though shown to be effective in many applications, can vary significantly in their response quality. In this paper, we investigate this problem of prompt fairness: specifically, the phrasing…

Bias Mitigation in Fine-tuning Pre-trained Models for Enhanced Fairness and Efficiency

2024-03-01 · Yixuan Zhang, Feng Zhou

Fine-tuning pre-trained models is a widely employed technique in numerous real-world applications. However, fine-tuning these models on new tasks can lead to unfair outcomes. This is due to the absence of generalization …

FairnessTransfer Learning