paper-with-me

Papers

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

2026-03-09 · Hongli Zhou, Hui Huang, Rui Zhang, Kehai Chen, Bing Xu, Conghui Zhu, Tiejun Zhao, Muyun Yang arxiv

Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by judgment biases. Accurately evaluating these biases is essential for ensuring the reliability of LLM-based judges. However, existing studies typically investigate limited biases under a single judge formulation, either generative or discriminative, lacking a comprehensive evaluation. To bridge this gap, we propose JudgeBiasBench, a benchmark for systematically quantifying biases in LLM-based judges. JudgeBiasBench defines a taxonomy of judgment biases across 4 dimensions, and constructs bias-augmented evaluation instances through a controlled bias injection pipeline, covering 12 representative bias types. We conduct extensive experiments across both generative and discriminative judges, revealing that current judges exhibit significant and diverse bias patterns that often compromise the reliability of automated evaluation. To mitigate judgment bias, we propose bias-aware training that explicitly incorporates bias-related attributes into the training process, encouraging judges to disentangle task-relevant quality from bias-correlated cues. By adopting reinforcement learning for generative judges and contrastive learning for discriminative judges, our methods effectively reduce judgment biases while largely preserving general evaluation capability.

📄 PDF Abstract BibTeX arXiv:2603.08091

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningContrastive Learning

Similar Papers 제목 키워드 기반

CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges

2024-10-20 · Haitao Li, Junjie Chen, Qingyao Ai, Zhumin Chu 외

The use of large language models (LLMs) as automated evaluation tools to assess the quality of generated natural language, known as LLMs-as-Judges, has demonstrated promising capabilities and is rapidly gaining widesprea…

FairnessPredictionSelection bias

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines

2026-04-25 · Sadman Kabir Soumik arxiv

LLM-as-a-Judge has become the dominant paradigm for evaluating language model outputs, yet LLM judges exhibit systematic biases that compromise evaluation reliability. We present a comprehensive empirical study comparing…

UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-Judge

2025-08-13 · Yang Zhang, Cunxiang Wang, Lindong Wu, Wenbo Yu 외 arxiv

Pairwise evaluation of Large Language Models (LLMs) is a common paradigm, but it is prone to preference bias, where judges systematically favor certain outputs, such as their own. This bias leads to inconsistent and skew…

Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data

2024-10-17 · Florian E. Dorner, Vivian Y. Nastl, Moritz Hardt

High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an important research ambition.…

It Is Different When Items Are Older: Debiasing Recommendations When Selection Bias and User Preferences Are Dynamic

2021-11-24 · Jin Huang, Harrie Oosterhuis, Maarten de Rijke

User interactions with recommender systems (RSs) are affected by user selection bias, e.g., users are more likely to rate popular items (popularity bias) or items that they expect to enjoy beforehand (positivity bias). M…

Recommendation SystemsSelection bias