paper-with-me

Papers

TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation

2024-10-04 · Jonathan Cook, Tim Rocktäschel, Jakob Foerster, Dennis Aumiller, Alex Wang

Given the widespread adoption and usage of Large Language Models (LLMs), it is crucial to have flexible and interpretable evaluations of their instruction-following ability. Preference judgments between model outputs have become the de facto evaluation standard, despite distilling complex, multi-faceted preferences into a single ranking. Furthermore, as human annotation is slow and costly, LLMs are increasingly used to make these judgments, at the expense of reliability and interpretability. In this work, we propose TICK (Targeted Instruct-evaluation with ChecKlists), a fully automated, interpretable evaluation protocol that structures evaluations with LLM-generated, instruction-specific checklists. We first show that, given an instruction, LLMs can reliably produce high-quality, tailored evaluation checklists that decompose the instruction into a series of YES/NO questions. Each question asks whether a candidate response meets a specific requirement of the instruction. We demonstrate that using TICK leads to a significant increase (46.4% $\to$ 52.2%) in the frequency of exact agreements between LLM judgements and human preferences, as compared to having an LLM directly score an output. We then show that STICK (Self-TICK) can be used to improve generation quality across multiple benchmarks via self-refinement and Best-of-N selection. STICK self-refinement on LiveBench reasoning tasks leads to an absolute gain of $+$7.8%, whilst Best-of-N selection with STICK attains $+$6.3% absolute improvement on the real-world instruction dataset, WildBench. In light of this, structured, multi-faceted self-improvement is shown to be a promising way to further advance LLM capabilities. Finally, by providing LLM-generated checklists to human evaluators tasked with directly scoring LLM responses to WildBench instructions, we notably increase inter-annotator agreement (0.194 $\to$ 0.256).

📄 PDF Abstract BibTeX arXiv:2410.03608

Code (0)

등록된 구현이 없습니다.

Tasks

AllInstruction Following

Similar Papers 제목 키워드 기반

Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation

2022-11-17 · Aleksandar Savkov, Francesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera 외

Evaluating automatically generated text is generally hard due to the inherently subjective nature of many aspects of the output quality. This difficulty is compounded in automatic consultation note generation by differin…

From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes

2025-07-23 · Karen Zhou, John Giorgi, Pranav Mani, Peng Xu 외 arxiv

AI-generated clinical notes are increasingly used in healthcare, but evaluating their quality remains a challenge due to high subjectivity and limited scalability of expert review. Existing automated metrics often fail t…

Multilingual CheckList: Generation and Evaluation

2022-03-24 · Karthikeyan K, Shaily Bhatt, Pankaj Singh, Somak Aditya 외

Multilingual evaluation benchmarks usually contain limited high-resource languages and do not test models for specific linguistic capabilities. CheckList is a template-based evaluation approach that tests models for spec…

DiversityMachine Translation

Layout-Corrector: Alleviating Layout Sticking Phenomenon in Discrete Diffusion Model

2024-09-25 · Shoma Iwai, Atsuki Osanai, Shunsuke Kitada, Shinichiro Omachi

Layout generation is a task to synthesize a harmonious layout with elements characterized by attributes such as category, position, and size. Human designers experiment with the placement and modification of elements to …

Layout Generation

IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation

2025-11-02 · Bosi Wen, Yilin Niu, Cunxiang Wang, Pei Ke 외 arxiv

Instruction-following is a fundamental ability of Large Language Models (LLMs), requiring their generated outputs to follow multiple constraints imposed in input instructions. Numerous studies have attempted to enhance t…

Reinforcement Learning