paper-with-me

홈 › Papers

Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation

2026-06-06 · Katelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi, Zongwan Cao, Chenjun Xu, Bingbing Wen, Su Lin Blodgett, Lucy Lu Wang arxiv

Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-documented protocols -- details that are frequently missing in current practice. In this work, we conduct a large-scale analysis of human evaluation protocols for evaluating long-form generation tasks in *CL conference publications from 2023--2025, including a full manual review of 284 papers and LLM-assisted analysis for another 1.8k+ papers. We define a set of 20 reportable criteria related to reproducibility of human evaluation studies, and apply these criteria to systematically examine reporting norms and practices within the community. We find widespread under-reporting of important aspects of human evaluation study design, leading to ambiguity about what was measured and how, who contributed judgments, and how judgments should be interpreted. Based on these findings, we outline actionable recommendations to support more transparent and reproducible reporting in future research. Our analysis code and annotated dataset can be found at: https://github.com/larchlab/Illusions-of-the-Gold-Standard

📄 PDF Abstract BibTeX arXiv:2606.07936

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

IllusionBench: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models

2025-01-01 · Yiming Zhang, ZiCheng Zhang, Xinyi Wei, Xiaohong Liu 외

Current Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have bee…

HallucinationMultiple-choice

BRI3L: A Brightness Illusion Image Dataset for Identification and Localization of Regions of Illusory Perception

2024-02-07 · Aniket Roy, Anirban Roy, Soma Mitra, Kuntal Ghosh

Visual illusions play a significant role in understanding visual perception. Current methods in understanding and evaluating visual illusions are mostly deterministic filtering based approach and they evaluate on a handf…

Benchmarking

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

2026-02-02 · Wenjin Hou, Wei Liu, Han Hu, Xiaoxiao Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have shown remarkable proficiency on general-purpose vision-language benchmarks, reaching or even exceeding human-level performance. However, these evaluations typically rely on s…

Visual Reasoning

The Illusion-Illusion: Vision Language Models See Illusions Where There are None

2024-12-07 · Tomer Ullman

Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something "really is" and how something "appears to be…

DiagnosticPhilosophy

A neuro-mathematical model for size and context related illusions

2019-08-27

We provide here a mathematical model of size/context illusions, inspired by the functional architecture of the visual cortex. We first recall previous models of scale and orientation, in particular the one in (Sarti et a…