paper-with-me

홈 › Papers

How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

2026-04-27 · Xinran Zhang arxiv

Safety benchmarks such as HarmBench rely on LLM judges to classify model responses as harmful or safe, yet the judge configuration, namely the combination of judge model and judge prompt, is typically treated as a fixed implementation detail. We show this assumption is problematic. Using a 2 x 2 x 3 factorial design, we construct 12 judge prompt variants along two axes, evaluation structure and instruction framing, and apply them using a single judge model, Claude Sonnet 4-6, producing 28,812 judgments over six target models and 400 HarmBench behaviors. We find that prompt wording alone, holding the judge model fixed, shifts measured harmful-response rates by up to 24.2 percentage points, with even within-condition surface rewording causing swings of up to 20.1 percentage points. Model safety rankings are moderately unstable, with mean Kendall tau = 0.89, and category-level sensitivity ranges from 39.6 percentage points for copyright to 0 percentage points for harassment. A supplementary multi-judge experiment using three judge models shows that judge-model choice adds further variance. Our results demonstrate that judge prompt wording is a substantial, previously under-examined source of measurement variance in safety benchmarking.

📄 PDF Abstract BibTeX arXiv:2604.24074

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks

2026-05-25 · Yanhang Li, Zhichao Fan, Zexin Zhuang arxiv

Pairwise model comparisons drawn from foundation-model benchmarks ("A is safer than B") are read as quantitative verdicts but hinge on harness choices benchmark papers under-specify. We close one theory-benchmark loop on…

Judge Reliability Harness: Stress Testing the Reliability of LLM Judges

2026-03-05 · Sunishchal Dev, Andrew Sloan, Joshua Kavner, Nicholas Kong 외 arxiv

We present the Judge Reliability Harness, an open source library for constructing validation suites that test the reliability of LLM judges. As LLM based scoring is widely deployed in AI benchmarks, more tooling is neede…

FreoStream:Enhancing Stream Guardrails via Future-Aware Reasoning and Safety-Aligned Optimization

2026-06-11 · Jianwei Wang, Guoyang Shen, Yanhong Wu, Haoran Li 외 arxiv

Stream guardrails enable token-level safety detection before full responses are generated. However, they often make overly conservative judgements and block those sensitive but safe tokens, which is known as over-refusal…

Configurable Reward Model for Balanced Safety Alignment

2026-05-28 · Zhengping Jiang, Mehran Khodabandeh, Akash Bharadwaj, Manik Bhandari 외 arxiv

Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned LLMs and standalone safety classifiers often fail to generalize to …

Data Augmentation

Cross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMs

2026-05-30 · Subhadip Mitra arxiv

Safety alignment in LLMs does not improve monotonically across model generations. Studying four generations of Google's Gemma family (7B-31B) with quality-diversity evolution (MAP-Elites) as an automated red-teaming prob…