paper-with-me

Papers

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

2025-08-25 · Khaoula Chehbouni, Mohammed Haddou, Jackie Chi Kit Cheung, Golnoosh Farnadi arxiv

Evaluating natural language generation (NLG) systems remains a core challenge of natural language processing (NLP), further complicated by the rise of large language models (LLMs) that aims to be general-purpose. Recently, large language models as judges (LLJs) have emerged as a promising alternative to traditional metrics, but their validity remains underexplored. This position paper argues that the current enthusiasm around LLJs may be premature, as their adoption has outpaced rigorous scrutiny of their reliability and validity as evaluators. Drawing on measurement theory from the social sciences, we identify and critically assess four core assumptions underlying the use of LLJs: their ability to act as proxies for human judgment, their capabilities as evaluators, their scalability, and their cost-effectiveness. We examine how each of these assumptions may be challenged by the inherent limitations of LLMs, LLJs, or current practices in NLG evaluation. To ground our analysis, we explore three applications of LLJs: text summarization, data annotation, and safety alignment. Finally, we highlight the need for more responsible evaluation practices in LLJs evaluation, to ensure that their growing role in the field supports, rather than undermines, progress in NLG.

📄 PDF Abstract BibTeX arXiv:2508.18076

Code (0)

등록된 구현이 없습니다.

Tasks

Text Summarization

Similar Papers 제목 키워드 기반

UXBench: Measuring the Actionability of LLM-Generated UX Critiques

2026-06-15 · Wenjie Wang, Yue Huang, Zipeng Ling, Han Bao 외 arxiv

Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repairs. Yet no controlled benchmark measures whether the resulting critiques are reli…

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models

2026-04-24 · Juan Manuel Contreras arxiv

Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models behave. Is this gap an artifact of forcing human trait categories onto LLMs, or someth…

Humans or LLMs as the Judge? A Study on Judgement Biases

2024-02-16 · Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang 외

Adopting human and large language models (LLM) as judges (a.k.a human- and LLM-as-a-judge) for evaluating the performance of LLMs has recently gained attention. Nonetheless, this approach concurrently introduces potentia…

Misinformation

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

2026-05-28 · Guneet Kohli arxiv

LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quan…

Natural Language Inference

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

2024-06-26 · Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott 외

There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case of proprietary mode…