paper-with-me

홈 › Papers

Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models

2026-03-23 · Tom Biskupski, Stephan Kleber arxiv

A Large Language Model (LLM) as judge evaluates the quality of victim Machine Learning (ML) models, specifically LLMs, by analyzing their outputs. An LLM as judge is the combination of one model and one specifically engineered judge prompt that contains the criteria for the analysis. The resulting automation of the analysis scales up the complex evaluation of the victim models' free-form text outputs by faster and more consistent judgments compared to human reviewers. Thus, quality and security assessments of LLMs can cover a wide range of the victim models' use cases. Being a comparably new technique, LLMs as judges lack a thorough investigation for their reliability and agreement to human judgment. Our work evaluates the applicability of LLMs as automated quality assessors of victim LLMs. We test the efficacy of 37 differently sized conversational LLMs in combination with 5 different judge prompts, the concept of a second-level judge, and 5 models fine-tuned for the task as assessors. As assessment objective, we curate datasets for eight different categories of judgment tasks and the corresponding ground-truth labels based on human assessments. Our empirical results show a high correlation of LLMs as judges with human assessments, when combined with a suitable prompt, in particular for GPT-4o, several open-source models with $\geqslant$ 32B parameters, and a few smaller models like Qwen2.5 14B.

📄 PDF Abstract BibTeX arXiv:2603.22214

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation

2024-01-10 · Zhaokun Jiang, Qianxi Lv, Ziyin Zhang, Lei Lei

Large language models have demonstrated parallel and even superior translation performance compared to neural machine translation (NMT) systems. However, existing comparative studies between them mainly rely on automated…

Machine TranslationNMTPrompt EngineeringTranslation

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

2026-03-09 · Hongli Zhou, Hui Huang, Rui Zhang, Kehai Chen 외 arxiv

Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by judgment biases. Accurately evaluating these biases is essential for ensu…

Reinforcement LearningContrastive Learning

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

2026-05-28 · Tianzhuo Yang, Zihan Shen, Zirui Mi, Zhaoyi Zhang 외 arxiv

Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. E…

Bias Detection

The Effect of Document Summarization on LLM-Based Relevance Judgments

2025-12-05 · Samaneh Mohtadi, Kevin Roitero, Stefano Mizzaro, Gianluca Demartini arxiv

Relevance judgments are central to the evaluation of Information Retrieval (IR) systems, but obtaining them from human annotators is costly and time-consuming. Large Language Models (LLMs) have recently been proposed as …

Document SummarizationInformation RetrievalText Summarization

SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics

2026-01-14 · Yunqiao Yang, Wenbo Li, Houxing Ren, Zimu Lu 외 arxiv

The rapid evolution of Large Language Models (LLMs) has fostered diverse paradigms for automated slide generation, ranging from code-driven layouts to image-centric synthesis. However, evaluating these heterogeneous syst…