paper-with-me

홈 › Papers

Mitigating Bias Using Model-Agnostic Data Attribution

2024-05-08 · Sander De Coninck, Wei-Cheng Wang, Sam Leroux, Pieter Simoens

Mitigating bias in machine learning models is a critical endeavor for ensuring fairness and equity. In this paper, we propose a novel approach to address bias by leveraging pixel image attributions to identify and regularize regions of images containing significant information about bias attributes. Our method utilizes a model-agnostic approach to extract pixel attributions by employing a convolutional neural network (CNN) classifier trained on small image patches. By training the classifier to predict a property of the entire image using only a single patch, we achieve region-based attributions that provide insights into the distribution of important information across the image. We propose utilizing these attributions to introduce targeted noise into datasets with confounding attributes that bias the data, thereby constraining neural networks from learning these biases and emphasizing the primary attributes. Our approach demonstrates its efficacy in enabling the training of unbiased classifiers on heavily biased datasets.

📄 PDF Abstract BibTeX arXiv:2405.05031

Code (0)

등록된 구현이 없습니다.

Tasks

Fairnessmodel

Similar Papers 제목 키워드 기반

Incorporating Priors with Feature Attribution on Text Classification

2019-06-19 · ACL 2019 7 · Frederick Liu, Besim Avci

Feature attribution methods, proposed recently, help users interpret the predictions of complex models. Our approach integrates feature attributions into the objective function to allow machine learning practitioners to …

ClassificationGeneral Classificationtext-classificationText Classification

Concept-Level Explainability for Auditing & Steering LLM Responses

2025-05-12 · Kenza Amara, Rita Sevastjanova, Mennatallah El-Assady

As large language models (LLMs) become widely deployed, concerns about their safety and alignment grow. An approach to steer LLM behavior, such as mitigating biases or defending against jailbreaks, is to identify which p…

Prompt EngineeringSemantic SimilaritySemantic Textual SimilarityText Generation

AIM: Attributing, Interpreting, Mitigating Data Unfairness

2024-06-13 · Zhining Liu, Ruizhong Qiu, Zhichen Zeng, Yada Zhu 외

Data collected in the real world often encapsulates historical discrimination against disadvantaged groups and individuals. Existing fair machine learning (FairML) research has predominantly focused on mitigating discrim…

Fairness

Fast Axiomatic Attribution for Neural Networks

2021-11-15 · NeurIPS 2021 12 · Robin Hesse, Simone Schaub-Meyer, Stefan Roth

Mitigating the dependence on spurious correlations present in the training dataset is a quickly emerging and important topic of deep learning. Recent approaches include priors on the feature attribution of a deep neural …

Guide the Learner: Controlling Product of Experts Debiasing Method Based on Token Attribution Similarities

2023-02-06 · Ali Modarressi, Hossein Amirkhani, Mohammad Taher Pilehvar

Several proposals have been put forward in recent years for improving out-of-distribution (OOD) performance through mitigating dataset biases. A popular workaround is to train a robust model by re-weighting training exam…

Decision MakingFact VerificationNatural Language Inference