paper-with-me

Papers

Check-Eval: A Checklist-based Approach for Evaluating Text Quality

2024-07-19 · Jayr Pereira, Andre Assumpcao, Roberto Lotufo

Evaluating the quality of text generated by large language models (LLMs) remains a significant challenge. Traditional metrics often fail to align well with human judgments, particularly in tasks requiring creativity and nuance. In this paper, we propose \textsc{Check-Eval}, a novel evaluation framework leveraging LLMs to assess the quality of generated text through a checklist-based approach. \textsc{Check-Eval} can be employed as both a reference-free and reference-dependent evaluation method, providing a structured and interpretable assessment of text quality. The framework consists of two main stages: checklist generation and checklist evaluation. We validate \textsc{Check-Eval} on two benchmark datasets: Portuguese Legal Semantic Textual Similarity and \textsc{SummEval}. Our results demonstrate that \textsc{Check-Eval} achieves higher correlations with human judgments compared to existing metrics, such as \textsc{G-Eval} and \textsc{GPTScore}, underscoring its potential as a more reliable and effective evaluation framework for natural language generation tasks. The code for our experiments is available at \url{https://anonymous.4open.science/r/check-eval-0DB4}

📄 PDF Abstract BibTeX arXiv:2407.14467

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Textual SimilarityText Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes

2025-07-23 · Karen Zhou, John Giorgi, Pranav Mani, Peng Xu 외 arxiv

AI-generated clinical notes are increasingly used in healthcare, but evaluating their quality remains a challenge due to high subjectivity and limited scalability of expert review. Existing automated metrics often fail t…

A Dynamic, Interpreted CheckList for Meaning-oriented NLG Metric Evaluation – through the Lens of Semantic Similarity Rating

2022-07-01 · *SEM (NAACL) 2022 7 · Laura Zeidler, Juri Opitz, Anette Frank

Evaluating the quality of generated text is difficult, since traditional NLG evaluation metrics, focusing more on surface form than meaning, often fail to assign appropriate scores.This is especially problematic for AMR-…

nlg evaluationSemantic SimilaritySemantic Textual Similarity

A Dynamic, Interpreted CheckList for Meaning-oriented NLG Metric Evaluation -- through the Lens of Semantic Similarity Rating

2022-05-24 · Laura Zeidler, Juri Opitz, Anette Frank

Evaluating the quality of generated text is difficult, since traditional NLG evaluation metrics, focusing more on surface form than meaning, often fail to assign appropriate scores. This is especially problematic for AMR…

nlg evaluationSemantic SimilaritySemantic Textual Similarity

Pitfalls of Conversational LLMs on News Debiasing

2024-04-09 · Ipek Baris Schlicht, Defne Altiok, Maryanne Taouk, Lucie Flek

This paper addresses debiasing in news editing and evaluates the effectiveness of conversational Large Language Models in this task. We designed an evaluation checklist tailored to news editors' perspectives, obtained ge…

Misinformation

Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation

2022-11-17 · Aleksandar Savkov, Francesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera 외

Evaluating automatically generated text is generally hard due to the inherently subjective nature of many aspects of the output quality. This difficulty is compounded in automatic consultation note generation by differin…