Do the Right Thing, Just Debias! Multi-Category Bias Mitigation Using LLMs
This paper tackles the challenge of building robust and generalizable bias mitigation models for language. Recognizing the limitations of existing datasets, we introduce ANUBIS, a novel dataset with 1507 carefully curated sentence pairs encompassing nine social bias categories. We evaluate state-of-the-art models like T5, utilizing Supervised Fine-Tuning (SFT), Reinforcement Learning (PPO, DPO), and In-Context Learning (ICL) for effective bias mitigation. Our analysis focuses on multi-class social bias reduction, cross-dataset generalizability, and environmental impact of the trained models. ANUBIS and our findings offer valuable resources for building more equitable AI systems and contribute to the development of responsible and unbiased technologies with broad societal impact.
Code (0)
등록된 구현이 없습니다.
Tasks
In-Context LearningSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Language-Assisted Debiasing and Smoothing for Foundation Model-Based Semi-Supervised Learning
Recent studies have focused on introducing pre-trained foundation models into semi-supervised learning (SSL) tasks. Nevertheless, these foundation models can exhibit biases toward different classes and tend to genera…
Pseudo LabelUnleashing the Potential of Model Bias for Generalized Category Discovery
Generalized Category Discovery is a significant and complex task that aims to identify both known and undefined novel categories from a set of unlabeled data, leveraging another labeled dataset containing only known cate…
DIG-FACE: De-biased Learning for Generalized Facial Expression Category Discovery
We introduce a novel task, Generalized Facial Expression Category Discovery (G-FACE), that discovers new, unseen facial expressions while recognizing known categories effectively. Even though there are generalized catego…
Facial Expression RecognitionTripletSpectrum-Aware Debiasing: A Modern Inference Framework with Applications to Principal Components Regression
Debiasing is a fundamental concept in high-dimensional statistics. While degrees-of-freedom adjustment is the state-of-the-art technique in high-dimensional linear regression, it is limited to i.i.d. samples and sub-Gaus…
compressed sensingregressionDCMT: A Direct Entire-Space Causal Multi-Task Framework for Post-Click Conversion Estimation
In recommendation scenarios, there are two long-standing challenges, i.e., selection bias and data sparsity, which lead to a significant drop in prediction accuracy for both Click-Through Rate (CTR) and post-click Conver…
counterfactualMulti-Task LearningSelection bias