paper-with-me

홈 › Papers

Large Language Model Bias Mitigation from the Perspective of Knowledge Editing

2024-05-15 · Ruizhe Chen, Yichen Li, Zikai Xiao, Zuozhu Liu

Existing debiasing methods inevitably make unreasonable or undesired predictions as they are designated and evaluated to achieve parity across different social groups but leave aside individual facts, resulting in modified existing knowledge. In this paper, we first establish a new bias mitigation benchmark BiasKE leveraging existing and additional constructed datasets, which systematically assesses debiasing performance by complementary metrics on fairness, specificity, and generalization. Meanwhile, we propose a novel debiasing method, Fairness Stamp (FAST), which enables editable fairness through fine-grained calibration on individual biased knowledge. Comprehensive experiments demonstrate that FAST surpasses state-of-the-art baselines with remarkable debiasing performance while not hampering overall model capability for knowledge preservation, highlighting the prospect of fine-grained debiasing strategies for editable fairness in LLMs.

📄 PDF Abstract BibTeX arXiv:2405.09341

Code (0)

등록된 구현이 없습니다.

Tasks

Fairnessknowledge editingLanguage ModelingLanguage ModellingLarge Language ModelSpecificity

Similar Papers 제목 키워드 기반

Multilingual Bias Detection and Mitigation for Indian Languages

2023-12-23 · Ankita Maity, Anubhav Sharma, Rudra Dhar, Tushar Abhishek 외

Lack of diverse perspectives causes neutrality bias in Wikipedia content leading to millions of worldwide readers getting exposed by potentially inaccurate information. Hence, neutrality bias detection and mitigation is …

Bias DetectionBinary ClassificationStyle Transfer

MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion

2025-07-03 · Xin Guan, PeiHsin Lin, Zekun Wu, Ze Wang 외 arxiv

Multiperspective Fusion (MPF) is a novel posttraining alignment framework for large language models (LLMs) developed in response to the growing need for easy bias mitigation. Built on top of the SAGED pipeline, an automa…

Prompt Engineering

CORGI-PM: A Chinese Corpus For Gender Bias Probing and Mitigation

2023-01-01 · Ge Zhang, Yizhi Li, Yaoyao Wu, Linyuan Zhang 외

As natural language processing (NLP) for gender bias becomes a significant interdisciplinary topic, the prevalent data-driven techniques such as large-scale language models suffer from data inadequacy and biased corpus, …

Sentence

Editable Fairness: Fine-Grained Bias Mitigation in Language Models

2024-08-07 · Ruizhe Chen, Yichen Li, Jianfei Yang, Joey Tianyi Zhou 외

Generating fair and accurate predictions plays a pivotal role in deploying large language models (LLMs) in the real world. However, existing debiasing methods inevitably generate unfair or incorrect predictions as they a…

Fairness

Simulating a Bias Mitigation Scenario in Large Language Models

2025-09-17 · Kiana Kiashemshaki, Mohammad Jalili Torkamani, Negin Mahmoudi, Meysam Shirdel Bilehsavar arxiv

Large Language Models (LLMs) have fundamentally transformed the field of natural language processing; however, their vulnerability to biases presents a notable obstacle that threatens both fairness and trust. This review…