paper-with-me

홈 › Papers

JudgeLRM: Large Reasoning Models as a Judge

2025-03-31 · Nuo Chen, Zhiyuan Hu, Qingyun Zou, Jiaying Wu, Qian Wang, Bryan Hooi, Bingsheng He

The rise of Large Language Models (LLMs) as evaluators offers a scalable alternative to human annotation, yet existing Supervised Fine-Tuning (SFT) for judges approaches often fall short in domains requiring complex reasoning. In this work, we investigate whether LLM judges truly benefit from enhanced reasoning capabilities. Through a detailed analysis of reasoning requirements across evaluation tasks, we reveal a negative correlation between SFT performance gains and the proportion of reasoning-demanding samples - highlighting the limitations of SFT in such scenarios. To address this, we introduce JudgeLRM, a family of judgment-oriented LLMs trained using reinforcement learning (RL) with judge-wise, outcome-driven rewards. JudgeLRM models consistently outperform both SFT-tuned and state-of-the-art reasoning models. Notably, JudgeLRM-3B surpasses GPT-4, and JudgeLRM-7B outperforms DeepSeek-R1 by 2.79% in F1 score, particularly excelling in judge tasks requiring deep reasoning.

📄 PDF Abstract BibTeX arXiv:2504.00050

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Adam 설명 없음
Multi-Head Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization

2025-05-19 · Austin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong 외

To keep pace with the increasing pace of large language models (LLM) development, model output evaluation has transitioned away from time-consuming human evaluation to automatic evaluation, where LLMs themselves are task…

Reinforcement Learning (RL)

MR. Judge: Multimodal Reasoner as a Judge

2025-05-19 · Renjie Pi, Felix Bai, Qibin Chen, Simon Wang 외

The paradigm of using Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) as evaluative judges has emerged as an effective approach in RLHF and inference-time scaling. In this work, we propose Multi…

MM-VetMultiple-choice

How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality

2026-04-08 · Minzhu Tu, Shiyu Ni, Keping Bi arxiv

Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One possible reason is that these judges lack s…

Mathematical ReasoningQuestion Answering

Prejudge-Before-Think: Enhancing Large Language Models at Test-Time by Process Prejudge Reasoning

2025-04-18 · Jianing Wang, Jin Jiang, Yang Liu, Mengdi Zhang 외

In this paper, we introduce a new \emph{process prejudge} strategy in LLM reasoning to demonstrate that bootstrapping with process prejudge allows the LLM to adaptively anticipate the errors encountered when advancing th…

Reinforcement Learning (RL)

Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning

2025-10-27 · Ran Xu, Jingjing Chen, Jiayu Ye, Yu Wu 외 arxiv

Large Language Models (LLMs) are widely used as judges to evaluate response quality, providing a scalable alternative to human evaluation. However, most LLM judges operate solely on intrinsic text-based reasoning, limiti…

Reinforcement Learning