paper-with-me

홈 › Papers

MABR: Multilayer Adversarial Bias Removal Without Prior Bias Knowledge

2024-08-10 · Maxwell J. Yin, Boyu Wang, Charles Ling

Models trained on real-world data often mirror and exacerbate existing social biases. Traditional methods for mitigating these biases typically require prior knowledge of the specific biases to be addressed, such as gender or racial biases, and the social groups associated with each instance. In this paper, we introduce a novel adversarial training strategy that operates independently of prior bias-type knowledge and protected attribute labels. Our approach proactively identifies biases during model training by utilizing auxiliary models, which are trained concurrently by predicting the performance of the main model without relying on task labels. Additionally, we implement these auxiliary models at various levels of the feature maps of the main model, enabling the detection of a broader and more nuanced range of bias features. Through experiments on racial and gender biases in sentiment and occupation classification tasks, our method effectively reduces social biases without the need for demographic annotations. Moreover, our approach not only matches but often surpasses the efficacy of methods that require detailed demographic insights, marking a significant advancement in bias mitigation techniques.

📄 PDF Abstract BibTeX arXiv:2408.05497

Code (1)

maxwellyin/mabr 공식 구현 pytorch

Tasks

Attribute

Similar Papers 제목 키워드 기반

Robsut Wrod Reocginiton via semi-Character Recurrent Neural Network

2016-08-07 · Keisuke Sakaguchi, Kevin Duh, Matt Post, Benjamin Van Durme

Language processing mechanism by humans is generally more robust than computers. The Cmabrigde Uinervtisy (Cambridge University) effect from the psycholinguistics literature has demonstrated such a robust word processing…

Spelling Correction

Multi-Agent Broad Reinforcement Learning for Intelligent Traffic Light Control

2022-03-08 · Ruijie Zhu, Lulu Li, Shuning Wu, Pei Lv 외

Intelligent Traffic Light Control System (ITLCS) is a typical Multi-Agent System (MAS), which comprises multiple roads and traffic lights.Constructing a model of MAS for ITLCS is the basis to alleviate traffic congestion…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

On Adversarial Removal of Hypothesis-only Bias in Natural Language Inference

2019-07-09 · SEMEVAL 2019 6 · Yonatan Belinkov, Adam Poliak, Stuart M. Shieber, Benjamin Van Durme 외

Popular Natural Language Inference (NLI) datasets have been shown to be tainted by hypothesis-only biases. Adversarial learning may help models ignore sensitive biases and spurious correlations in data. We evaluate wheth…

Natural Language Inference

Using Adversarial Debiasing to Remove Bias from Word Embeddings

2021-07-21 · Dana Kenna

Word Embeddings have been shown to contain the societal biases present in the original corpora. Existing methods to deal with this problem have been shown to only remove superficial biases. The method of Adversarial Debi…

Word Embeddings

SABAF: Removing Strong Attribute Bias from Neural Networks with Adversarial Filtering

2023-11-13 · Jiazhi Li, Mahyar Khayatkhoei, Jiageng Zhu, Hanchen Xie 외

Ensuring a neural network is not relying on protected attributes (e.g., race, sex, age) for prediction is crucial in advancing fair and trustworthy AI. While several promising methods for removing attribute bias in neura…

Attribute