paper-with-me

Papers

Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?

2025-08-21 · Momoka Furuhashi, Kouta Nakayama, Takashi Kodama, Saku Sugawara arxiv

Automatic evaluation of generative tasks using large language models faces challenges due to ambiguous criteria. Although automatic checklist generation is a potentially promising approach, its usefulness remains underexplored. We investigate whether checklists should be used for all questions or selectively, generate them using six methods, evaluate their effectiveness across eight model sizes, and identify checklist items that correlate with human evaluations. Through experiments on pairwise comparison and direct scoring tasks, we find that selective checklist use tends to improve evaluation performance in pairwise settings, while its benefits are less consistent in direct scoring. Our analysis also shows that even checklist items with low correlation to human scores often reflect human-written criteria, indicating potential inconsistencies in human evaluation. These findings highlight the need to more clearly define objective evaluation criteria to guide both human and automatic evaluations. \footnote{Our code is available at~https://github.com/momo0817/checklist-effectiveness-study

📄 PDF Abstract BibTeX arXiv:2508.15218

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multilingual CheckList: Generation and Evaluation

2022-03-24 · Karthikeyan K, Shaily Bhatt, Pankaj Singh, Somak Aditya 외

Multilingual evaluation benchmarks usually contain limited high-resource languages and do not test models for specific linguistic capabilities. CheckList is a template-based evaluation approach that tests models for spec…

DiversityMachine Translation

Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation

2022-11-17 · Aleksandar Savkov, Francesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera 외

Evaluating automatically generated text is generally hard due to the inherently subjective nature of many aspects of the output quality. This difficulty is compounded in automatic consultation note generation by differin…

Perturbation CheckLists for Evaluating NLG Evaluation Metrics

2021-09-13 · EMNLP 2021 11 · Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan 외

Natural Language Generation (NLG) evaluation is a multifaceted task requiring assessment of multiple desirable criteria, e.g., fluency, coherency, coverage, relevance, adequacy, overall quality, etc. Across existing data…

Data-to-Text Generationnlg evaluationText Generation

Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist

2024-07-11 · ZiHao Zhou, Shudong Liu, Maizhen Ning, Wei Liu 외

Exceptional mathematical reasoning ability is one of the key features that demonstrate the power of large language models (LLMs). How to comprehensively define and evaluate the mathematical abilities of LLMs, and even re…

GSM8KMathMathematical Reasoning

Learning Predictive Checklists with Probabilistic Logic Programming

2024-11-25 · Yukti Makhija, Edward De Brouwer, Rahul G. Krishnan

Checklists have been widely recognized as effective tools for completing complex tasks in a systematic manner. Although originally intended for use in procedural tasks, their interpretability and ease of use have led to …

Time Series