Validating LLM-as-a-Judge Systems in the Absence of Gold Labels
The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, has come to play a critical role in scaling and standardizing GenAI evaluations. To validate judge systems, evaluators collect multiple human ratings for each item in a validation corpus, and then aggregate the ratings into a single, per-item gold label rating. High agreement rates between these gold labels and judge system ratings are then taken as a sign of good judge system performance. In many cases, however, items or rating criteria may be ambiguous, or there may be principled disagreement among human raters. In such settings, gold labels may not exist for many of the items. In this paper, we introduce a framework for LLM-as-a-judge validation in the absence of gold labels. We present a theoretical analysis drawing connections between different measures of judge system performance under different rating elicitation and aggregation schemes. We also demonstrate empirically that existing validation approaches can select judge systems that are highly suboptimal, performing as much as 34% worse than the systems selected by alternative approaches that we describe. Based on our findings, we provide concrete recommendations for developing more reliable approaches to LLM-as-a-judge validation.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
Unsupervised methods are widely used to induce latent semantic structure from large text collections, yet their outputs often contain incoherent, redundant, or poorly grounded clusters that are difficult to validate with…
Representation LearningTopic ModelsA Finite-Calibration Regime Map for LLM Judge Panels
Deploying an LLM judge panel spends human labels on fitting a calibrator, constructing candidate judge paths, and validating which candidate to deploy. We study when finite labels should support a low-dimensional stacker…
Named Entity Recognition in the Legal Domain using a Pointer Generator Network
Named Entity Recognition (NER) is the task of identifying and classifying named entities in unstructured text. In the legal domain, named entities of interest may include the case parties, judges, names of courts, case n…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1GLaPE: Gold Label-agnostic Prompt Evaluation and Optimization for Large Language Model
Despite the rapid progress of large language models (LLMs), their task performance remains sensitive to prompt design. Recent studies have explored leveraging the LLM itself as an optimizer to identify optimal prompts th…
Language ModelingLanguage ModellingLarge Language ModelLegal Entity Extraction using a Pointer Generator Network
Named Entity Recognition (NER) is the task of identifying and classifying named entities in unstructured text. In the legal domain, named entities of interest may include the case parties, judges, names of courts, cas…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1