paper-with-me

Papers

Validating LLM-as-a-Judge Systems in the Absence of Gold Labels

2025-03-07 · Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach, Zhiwei Steven Wu, Alexandra Chouldechova

The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, has come to play a critical role in scaling and standardizing GenAI evaluations. To validate judge systems, evaluators collect multiple human ratings for each item in a validation corpus, and then aggregate the ratings into a single, per-item gold label rating. High agreement rates between these gold labels and judge system ratings are then taken as a sign of good judge system performance. In many cases, however, items or rating criteria may be ambiguous, or there may be principled disagreement among human raters. In such settings, gold labels may not exist for many of the items. In this paper, we introduce a framework for LLM-as-a-judge validation in the absence of gold labels. We present a theoretical analysis drawing connections between different measures of judge system performance under different rating elicitation and aggregation schemes. We also demonstrate empirically that existing validation approaches can select judge systems that are highly suboptimal, performing as much as 34% worse than the systems selected by alternative approaches that we describe. Based on our findings, we provide concrete recommendations for developing more reliable approaches to LLM-as-a-judge validation.

📄 PDF Abstract BibTeX arXiv:2503.05965

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs

2026-04-08 · Tunazzina Islam arxiv

Unsupervised methods are widely used to induce latent semantic structure from large text collections, yet their outputs often contain incoherent, redundant, or poorly grounded clusters that are difficult to validate with…

Representation LearningTopic Models

A Finite-Calibration Regime Map for LLM Judge Panels

2026-05-31 · Bin Zhu, Yanghui Rao arxiv

Deploying an LLM judge panel spends human labels on fitting a calibrator, constructing candidate judge paths, and validating which candidate to deploy. We study when finite labels should support a low-dimensional stacker…

Named Entity Recognition in the Legal Domain using a Pointer Generator Network

2020-12-17 · Stavroula Skylaki, Ali Oskooei, Omar Bari, Nadja Herger 외

Named Entity Recognition (NER) is the task of identifying and classifying named entities in unstructured text. In the legal domain, named entities of interest may include the case parties, judges, names of courts, case n…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

GLaPE: Gold Label-agnostic Prompt Evaluation and Optimization for Large Language Model

2024-02-04 · Xuanchang Zhang, Zhuosheng Zhang, Hai Zhao

Despite the rapid progress of large language models (LLMs), their task performance remains sensitive to prompt design. Recent studies have explored leveraging the LLM itself as an optimizer to identify optimal prompts th…

Language ModelingLanguage ModellingLarge Language Model

Legal Entity Extraction using a Pointer Generator Network

2022-01-20 · International Conference on Data Mining Workshops (ICDMW) 2022 1 · Stavroula Skylaki, Ali Oskooei, Omar Bari, Nadja Herger 외

Named Entity Recognition (NER) is the task of identifying and classifying named entities in unstructured text. In the legal domain, named entities of interest may include the case parties, judges, names of courts, cas…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1