paper-with-me

홈 › Papers

Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases

2026-01-07 · Hui Huang, Xuanxin Wu, Muyun Yang, Yuki Arase arxiv

This paper presents the first systematic comparison investigating whether Large Reasoning Models (LRMs) are superior judges to non-reasoning LLMs. Our empirical analysis yields four key findings: 1) LRMs outperform non-reasoning LLMs in terms of judgment accuracy, particularly on reasoning-intensive tasks; 2) LRMs demonstrate superior evaluation instruction-following capabilities; 3) LRMs exhibit enhanced robustness against adversarial attacks targeting judgment tasks; 4) However, LRMs still exhibit strong evaluation biases. To mitigate this bias vulnerability, we propose PlanJudge, a lightweight evaluation strategy that prompts the model to generate an explicit evaluation plan before executing the judgment. Despite its simplicity, our experiments demonstrate that PlanJudge significantly mitigates biases in LLM-as-a-Judge while preserving overall judgment accuracy.

📄 PDF Abstract BibTeX arXiv:2601.03630

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

2025-04-14 · Qian Wang, Zhanzhi Lou, Zhenheng Tang, Nuo Chen 외

Large Reasoning Models (LRMs) like DeepSeek-R1 and OpenAI-o1 have demonstrated remarkable reasoning capabilities, raising important questions about their biases in LLM-as-a-judge settings. We present a comprehensive benc…

In-Context LearningPosition

How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality

2026-04-08 · Minzhu Tu, Shiyu Ni, Keping Bi arxiv

Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One possible reason is that these judges lack s…

Mathematical ReasoningQuestion Answering

J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization

2025-05-19 · Austin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong 외

To keep pace with the increasing pace of large language models (LLM) development, model output evaluation has transitioned away from time-consuming human evaluation to automatic evaluation, where LLMs themselves are task…

Reinforcement Learning (RL)

Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation

2026-02-07 · Jiangnan Fang, Cheng-Tse Liu, Hanieh Deilamsalehy, Nesreen K. Ahmed 외 arxiv

Large language model (LLM) judges have often been used alongside traditional, algorithm-based metrics for tasks like summarization because they better capture semantic information, are better at reasoning, and are more r…

Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems

2025-10-14 · Jiaxin Gao, Chen Chen, Yanwen Jia, Xueluan Gong 외 arxiv

Large Language Models (LLMs) are increasingly being used to autonomously evaluate the quality of content in communication systems, e.g., to assess responses in telecom customer support chatbots. However, the impartiality…