paper-with-me

홈 › Papers

Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation

2019-01-02 · IJCNLP 2019 11 · Cristina Garbacea, Samuel Carton, Shiyan Yan, Qiaozhu Mei

We conduct a large-scale, systematic study to evaluate the existing evaluation methods for natural language generation in the context of generating online product reviews. We compare human-based evaluators with a variety of automated evaluation procedures, including discriminative evaluators that measure how well machine-generated text can be distinguished from human-written text, as well as word overlap metrics that assess how similar the generated text compares to human-written references. We determine to what extent these different evaluators agree on the ranking of a dozen of state-of-the-art generators for online product reviews. We find that human evaluators do not correlate well with discriminative evaluators, leaving a bigger question of whether adversarial accuracy is the correct objective for natural language generation. In general, distinguishing machine-generated text is challenging even for human evaluators, and human decisions correlate better with lexical overlaps. We find lexical diversity an intriguing metric that is indicative of the assessments of different evaluators. A post-experiment survey of participants provides insights into how to evaluate and improve the quality of natural language generation systems.

📄 PDF Abstract BibTeX arXiv:1901.00398

Code (1)

Crista23/JudgeTheJudges 공식 구현

Tasks

DiversityReview GenerationText Generation

Similar Papers 제목 키워드 기반

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

2026-08-27 · Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin 외 arxiv

Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to mode…

Model extraction

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

2026-08-27 · Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian, Sai Harshitha Aluru arxiv

LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systema…

A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

2026-05-29 · Yi Zhao, Siqi Wang, Zhe Hu, Yushi Li 외 arxiv

AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradigm may offer a promising alternative, although it has mostly been studied in gen…

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

2026-04-26 · Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang arxiv

Large language models are widely adopted as automated evaluation judges, yet the stability of their verdicts under semantically equivalent prompt rephrasings remains largely unexamined. We conduct a systematic empirical …

JuStRank: Benchmarking LLM Judges for System Ranking

2024-12-12 · Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim 외

Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such evaluations make the us…

Benchmarking