paper-with-me

홈 › Papers

On Cost-Effective LLM-as-a-Judge Improvement Techniques

2026-04-15 · Ryan Lail, Luke Markham arxiv

Using a language model to score or rank candidate responses has become a scalable alternative to human evaluation in reinforcement learning from human feedback (RLHF) pipelines, benchmarking, and application layer evaluations. However, output reliability depends heavily on prompting and aggregation strategy. We present an empirical investigation of four drop-in techniques -- ensemble scoring, task-specific criteria injection, calibration context, and adaptive model escalation -- for improving LLM judge accuracy on RewardBench 2, with a unifying lens of noise control on the stochastic judge: ensembling as Monte Carlo averaging over per-call noise, criteria injection as between-response discrimination sharpening, and per-response score variance as an uncertainty signal. Ensemble scoring and task-specific criteria injection (the latter virtually cost free) together reach up to 85.8% accuracy, +13.5pp over baseline. Calibration context and adaptive model escalation also improve over baseline but are dominated by criteria + ensembling on the cost-accuracy Pareto frontier. Small models benefit disproportionately from ensembling, making high-accuracy LLM judges accessible at low cost. We show that these techniques generalise across model providers, evaluating on both OpenAI GPT and Anthropic Claude families.

📄 PDF Abstract BibTeX arXiv:2604.13717

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

2024-11-26 · Shijian Deng, Wentian Zhao, Yu-Jhe Li, Kun Wan 외

Self-improvement in multimodal large language models (MLLMs) is crucial for enhancing their reliability and robustness. However, current methods often rely heavily on MLLMs themselves as judges, leading to high computati…

Hallucination

Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning

2025-08-03 · Derin Cayir, Renjie Tao, Rashi Rungta, Kai Sun 외 arxiv

Large Language Models (LLMs) have demonstrated remarkable progress through preference-based fine-tuning, which critically depends on the quality of the underlying training data. While human feedback is essential for impr…

A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

2026-05-29 · Yi Zhao, Siqi Wang, Zhe Hu, Yushi Li 외 arxiv

AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradigm may offer a promising alternative, although it has mostly been studied in gen…

Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

2024-07-28 · Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu 외

Large Language Models (LLMs) are rapidly surpassing human knowledge in many domains. While improving these models traditionally relies on costly human data, recent self-rewarding mechanisms (Yuan et al., 2024) have shown…

HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation

2026-01-27 · Kla Tantithamthavorn, Hong Yi Lin, Patanamon Thongtanunam, Wachiraphan Charoenwet 외 arxiv

Large Language models (LLMs) have shown strong capabilities in code review automation, such as review comment generation, yet they suffer from hallucinations -- where the generated review comments are ungrounded in the a…