paper-with-me

Papers

J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization

2025-05-19 · Austin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong, Shafiq Joty

To keep pace with the increasing pace of large language models (LLM) development, model output evaluation has transitioned away from time-consuming human evaluation to automatic evaluation, where LLMs themselves are tasked with assessing and critiquing other model outputs. LLM-as-judge models are a class of generative evaluators that excel in evaluating relatively simple domains, like chat quality, but struggle in reasoning intensive domains where model responses contain more substantive and challenging content. To remedy existing judge shortcomings, we explore training judges with reinforcement learning (RL). We make three key contributions: (1) We propose the Equivalent Initial State Group Relative Policy Optimization (EIS-GRPO) algorithm, which allows us to train our judge to be robust to positional biases that arise in more complex evaluation settings. (2) We introduce ReasoningJudgeBench, a benchmark that evaluates judges in diverse reasoning settings not covered by prior work. (3) We train Judge for Reasoning (J4R), a 7B judge trained with EIS-GRPO that outperforms GPT-4o and the next best small judge by 6.7% and 9%, matching or exceeding the performance of larger GRPO-trained judges on both JudgeBench and ReasoningJudgeBench.

📄 PDF Abstract BibTeX arXiv:2505.13346

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)

Similar Papers 제목 키워드 기반

R$^3$-SQL: Ranking Reward and Resampling for Text-to-SQL

2026-04-28 · Hojae Han, Yeonseok Jeong, Seung-won Hwang, Zhewei Yao 외 arxiv

Modern Text-to-SQL systems generate multiple candidate SQL queries and rank them to judge a final prediction. However, existing methods face two limitations. First, they often score functionally equivalent SQL queries in…

Does the Judge Prefer English? Evaluating Language-Switching Invariance in LLM-as-a-Judge

2026-06-12 · Shaojie Yin arxiv

Large language models (LLMs) are now widely used as automatic judges for open-ended instruction-following evaluation. This practice is convenient, scalable, and often more semantically aware than reference-based metrics,…

When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning

2026-03-22 · Zhengxian Wu, Kai Shi, Chuanrui Zhang, Zirui Liao 외 arxiv

Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-model distillation, both of which are co…

Mathematical ReasoningMultimodal Reasoning

Counterfactual equivalence for POMDPs, and underlying deterministic environments

2018-01-11 · Stuart Armstrong

Partially Observable Markov Decision Processes (POMDPs) are rich environments often used in machine learning. But the issue of information and causal structures in POMDPs has been relatively little studied. This paper pr…

BIG-bench Machine Learningcounterfactual

Identifying equivalents of specialized verbs in a bilingual comparable corpus of judgments: A frame-based methodology

2012-05-01 · LREC 2012 5 · Janine Pimentel

Multilingual terminological resources do not always include the equivalents of specialized verbs that occur in legal texts. This study aims to bridge that gap by proposing a methodology to assign the equivalents of this …