paper-with-me

홈 › Papers

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

2025-04-14 · Qian Wang, Zhanzhi Lou, Zhenheng Tang, Nuo Chen, Xuandong Zhao, Wenxuan Zhang, Dawn Song, Bingsheng He

Large Reasoning Models (LRMs) like DeepSeek-R1 and OpenAI-o1 have demonstrated remarkable reasoning capabilities, raising important questions about their biases in LLM-as-a-judge settings. We present a comprehensive benchmark comparing judging biases between LLMs and LRMs across both subjective preference-alignment datasets and objective fact-based datasets. Through investigation of bandwagon, authority, position, and distraction biases, we uncover four key findings: (1) despite their advanced reasoning capabilities, LRMs remain susceptible to the above biases; (2) LRMs demonstrate better robustness than LLMs specifically on fact-related datasets; (3) LRMs exhibit notable position bias, preferring options in later positions; and (4) we identify a novel "superficial reflection bias" where phrases mimicking reasoning (e.g., "wait, let me think...") significantly influence model judgments. To address these biases, we design and evaluate three mitigation strategies: specialized system prompts that reduce judging biases by up to 19\% in preference alignment datasets and 14\% in fact-related datasets, in-context learning that provides up to 27\% improvement on preference tasks but shows inconsistent results on factual tasks, and a self-reflection mechanism that reduces biases by up to 10\% in preference datasets and 16\% in fact-related datasets, with self-reflection proving particularly effective for LRMs. Our work provides crucial insights for developing more reliable LLM-as-a-Judge frameworks, especially as LRMs become increasingly deployed as automated judges.

📄 PDF Abstract BibTeX arXiv:2504.09946

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningPosition

Similar Papers 제목 키워드 기반

Evaluating Strategic Reasoning in Forecasting Agents

2026-04-28 · Tom Liptay, Dan Schwarz, Rafael Poyiadzi, Jack Wildman 외 arxiv

Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to the Future 2 (BTF-2), 1,417 pastcasting questions with a frozen 15M-d…

MuSciClaims: Multimodal Scientific Claim Verification

2025-06-05 · Yash Kumar Lal, Manikanta Bandham, Mohammad Saqib Hasan, Apoorva Kashi 외

Assessing scientific claims requires identifying, extracting, and reasoning with multimodal data expressed in information-rich figures in scientific literature. Despite the large body of work in scientific QA, figure cap…

ArticlesClaim VerificationDiagnosticMultimodal Reasoning

To Mask or to Mirror: Human-AI Alignment in Collective Reasoning

2025-10-02 · Crystal Qian, Aaron Parisi, Clémentine Bouleau, Vivian Tsai 외 arxiv

As large language models (LLMs) are increasingly used to model and augment collective decision-making, it is critical to examine their alignment with human social reasoning. We present an empirical framework for assessin…

CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

2025-05-19 · Jinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang 외

Recent Large Reasoning Models significantly improve the reasoning ability of Large Language Models by learning to reason, exhibiting the promising performance in solving complex tasks. LRMs solve tasks that require compl…

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

2025-07-14 · Hongchao Jiang, Yiming Chen, Yushi Cao, Hung-Yi Lee 외

Large Language Models (LLMs) have significantly advanced the state-of-the-art in various coding tasks. Beyond directly answering user queries, LLMs can also serve as judges, assessing and comparing the quality of respons…

BenchmarkingCode GenerationCode Repair