paper-with-me

홈 › Papers

Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models

2025-03-10 · Kefan Song, Jin Yao, Runnan Jiang, Rohan Chandra, Shangtong Zhang

As Large Language Models (LLMs) become increasingly powerful and accessible to human users, ensuring fairness across diverse demographic groups, i.e., group fairness, is a critical ethical concern. However, current fairness and bias research in LLMs is limited in two aspects. First, compared to traditional group fairness in machine learning classification, it requires that the non-sensitive attributes, in this case, the prompt questions, be the same across different groups. In many practical scenarios, different groups, however, may prefer different prompt questions and this requirement becomes impractical. Second, it evaluates group fairness only for the LLM's final output without identifying the source of possible bias. Namely, the bias in LLM's output can result from both the pretraining and the finetuning. For finetuning, the bias can result from both the RLHF procedure and the learned reward model. Arguably, evaluating the group fairness of each component in the LLM pipeline could help develop better methods to mitigate the possible bias. Recognizing those two limitations, this work benchmarks the group fairness of learned reward models. By using expert-written text from arXiv, we are able to benchmark the group fairness of reward models without requiring the same prompt questions across different demographic groups. Surprisingly, our results demonstrate that all the evaluated reward models (e.g., Nemotron-4-340B-Reward, ArmoRM-Llama3-8B-v0.1, and GRM-llama3-8B-sftreg) exhibit statistically significant group unfairness. We also observed that top-performing reward models (w.r.t. canonical performance metrics) tend to demonstrate better group fairness.

📄 PDF Abstract BibTeX arXiv:2503.07806

Code (0)

등록된 구현이 없습니다.

Tasks

AllBenchmarkingFairness

Similar Papers 제목 키워드 기반

FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods

2023-06-15 · Xiaotian Han, Jianfeng Chi, Yu Chen, Qifan Wang 외

This paper introduces the Fair Fairness Benchmark (\textsf{FFB}), a benchmarking framework for in-processing group fairness methods. Ensuring fairness in machine learning is important for ethical compliance. However, the…

BenchmarkingFairnessGPU

Software Fairness Dilemma: Is Bias Mitigation a Zero-Sum Game?

2025-08-05 · Zhenpeng Chen, Xinyue Li, Jie M. Zhang, Weisong Sun 외 arxiv

Fairness is a critical requirement for Machine Learning (ML) software, driving the development of numerous bias mitigation methods. Previous research has identified a leveling-down effect in bias mitigation for computer …

Social Bias Probing: Fairness Benchmarking for Language Models

2023-11-15 · Marta Marchiori Manerba, Karolina Stańczak, Riccardo Guidotti, Isabelle Augenstein

While the impact of social biases in language models has been recognized, prior methods for bias evaluation have been limited to binary association tests on small datasets, limiting our understanding of bias complexities…

BenchmarkingFairnessProbing Language Models

Fair-GPTQ: Bias-Aware Quantization for Large Language Models

2025-09-18 · Irina Proskurina, Guillaume Metzler, Julien Velcin arxiv

The high memory demands of generative language models have drawn attention to quantization, which reduces memory usage by mapping model weights to lower-precision integers. However, recent empirical studies show that, wh…

Text Generation

Benchmarking Gender and Political Bias in Large Language Models

2025-09-07 · Jinrui Yang, Xudong Han, Timothy Baldwin arxiv

We introduce EuroParlVote, a novel benchmark for evaluating large language models (LLMs) in politically sensitive contexts. It links European Parliament debate speeches to roll-call vote outcomes and includes rich demogr…