paper-with-me

홈 › Papers

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering

2026-04-18 · Zixiao Zhao, Amirreza Esmaeili, Fatemeh Fard arxiv

Large Language Models are increasingly used as judges to evaluate code artifacts when exhaustive human review or executable test coverage is unavailable. LLM-judge is increasingly relevant in agentic software engineering workflows, where it can help rank candidate solutions and guide patch selection. While attractive for scale, current practice lacks a principled account of reliability and bias: repeated evaluations of the same case can disagree; small prompt edits can swing outcomes; and seemingly semantics-preserving, human-equivalent perturbations may elicit divergent verdicts. This paper studies LLM-as-a-Judge for code through a measurement-first lens. We analyze two pointwise judging regimes across code generation, code repair task, and test generation, and we systematically probe prompt-induced biases. Our study considers difficulty levels for repeated runs and controlled prompt interventions that isolate one presentation cue at a time, and it evaluates judges using consistency and sensitivity to bias. We find that judge decisions are highly sensitive to prompt biases even when the underlying code snippet is unchanged. Across all three tasks, several biases systematically shift preferences toward the option favored by the prompt, improving accuracy when that option aligns with the gold answer but substantially reducing it otherwise. In some settings, these effects are large enough to change task-level conclusions and alter relative model rankings. These findings show that reported judge performance may reflect prompt artifacts rather than stable assessment ability, posing a direct threat to the validity and reproducibility of code evaluation. We therefore argue that LLM-as-a-Judge studies should report bias sensitivity alongside accuracy and incorporate explicit controls to support more trustworthy model comparison in software engineering.

📄 PDF Abstract BibTeX arXiv:2604.16790

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationCode Repair

Similar Papers 제목 키워드 기반

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning

2026-06-17 · Zilong Zhang, Yi-Ting Hung, Lei Ding, Chi-Kuang Yeh arxiv

Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic quality, most notably verbosity bias. Me…

Comparing Developer and LLM Biases in Code Evaluation

2026-03-25 · Aditya Mittal, Ryan Shar, Zichu Wu, Shyam Agarwal 외 arxiv

As LLMs are increasingly used as judges in code applications, they should be evaluated in realistic interactive settings that capture partial context and ambiguous intent. We present TRACE (Tool for Rubric Analysis in Co…

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

2026-06-18 · Zilong Zhang, Yi-Ting Hung, Weiyi He, Junxi Zhang 외 arxiv

Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their preferences remain imperfect proxies for human …

Auditing Reward Hackability in Code RL Training Environments

2026-06-14 · Shreshth Rajan arxiv

We measure the rate at which code RL environments accept incorrect solutions as correct. On a 49-task sample of SWE-bench Verified, 28.5% of tasks have test suites weak enough that a Docker-verified incorrect patch passe…

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

2026-07-09 · Zongyou Yang, Yinghan Hou, Xiaokun Yang arxiv

An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replacement ambiguity as a measurement-validity problem. Across four judgment…