paper-with-me

홈 › Papers

BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses

2025-09-30 · Xin Xu, Xunzhi He, Churan Zhi, Ruizhe Chen, Julian McAuley, Zexue He arxiv

Existing studies on bias mitigation methods for large language models (LLMs) use diverse baselines and metrics to evaluate debiasing performance, leading to inconsistent comparisons among them. Moreover, their evaluations are mostly based on the comparison between LLMs' probabilities of biased and unbiased contexts, which ignores the gap between such evaluations and real-world use cases where users interact with LLMs by reading model responses and expect fair and safe outputs rather than LLMs' probabilities. To enable consistent evaluation across debiasing methods and bridge this gap, we introduce BiasFreeBench, an empirical benchmark that comprehensively compares eight mainstream bias mitigation techniques (covering four prompting-based and four training-based methods) on two test scenarios (multi-choice QA and open-ended multi-turn QA) by reorganizing existing datasets into a unified query-response setting. We further introduce a response-level metric, Bias-Free Score, to measure the extent to which LLM responses are fair, safe, and anti-stereotypical. Debiasing performances are systematically compared and analyzed across key dimensions: the prompting vs. training paradigm, model size, and generalization of different training strategies to unseen bias types. We release our benchmark, aiming to establish a unified testbed for bias mitigation research.

📄 PDF Abstract BibTeX arXiv:2510.00232

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Co$^2$PT: Mitigating Bias in Pre-trained Language Models through Counterfactual Contrastive Prompt Tuning

2023-10-19 · Xiangjue Dong, Ziwei Zhu, Zhuoer Wang, Maria Teleki 외

Pre-trained Language Models are widely used in many important real-world applications. However, recent studies show that these models can encode social biases from large pre-training corpora and even amplify biases in do…

counterfactual

Towards Understanding and Mitigating Social Biases in Language Models

2021-06-24 · Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan Salakhutdinov

As machine learning methods are deployed in real-world settings such as healthcare, legal systems, and social science, it is crucial to recognize how they shape social biases and stereotypes in these sensitive decision-m…

Decision MakingFairnessText Generation

On Evaluating and Mitigating Gender Biases in Multilingual Settings

2023-07-04 · Aniket Vashishtha, Kabir Ahuja, Sunayana Sitaram

While understanding and removing gender biases in language models has been a long-standing problem in Natural Language Processing, prior research work has primarily been limited to English. In this work, we investigate s…

VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs

2025-02-23 · Yiming Yang, Yangyang Guo, Hui Lu, Yan Wang

Recently, Large Vision-Language Models (LVLMs) have made significant strides across diverse multimodal tasks and benchmarks. This paper reveals a largely under-explored problem from existing video-involved LVLMs - langua…

Benchmarking

Locating and Mitigating Gender Bias in Large Language Models

2024-03-21 · Yuchen Cai, Ding Cao, Rongxi Guo, Yaqin Wen 외

Large language models(LLM) are pre-trained on extensive corpora to learn facts and human cognition which contain human preferences. However, this process can inadvertently lead to these models acquiring biases and stereo…

knowledge editingLanguage ModellingLarge Language ModelSentence