paper-with-me

홈 › Papers

Human vs Automatic Metrics: on the Importance of Correlation Design

2018-05-29 · Anastasia Shimorina

This paper discusses two existing approaches to the correlation analysis between automatic evaluation metrics and human scores in the area of natural language generation. Our experiments show that depending on the usage of a system- or sentence-level correlation analysis, correlation results between automatic scores and human judgments are inconsistent.

📄 PDF Abstract BibTeX arXiv:1805.11474

Code (1)

https://gitlab.com/webnlg/webnlg-human-evaluation 공식 구현

Tasks

SentenceText Generation

Similar Papers 제목 키워드 기반

LCEval: Learned Composite Metric for Caption Evaluation

2020-12-24 · Naeha Sharif, Lyndon White, Mohammed Bennamoun, Wei Liu 외

Automatic evaluation metrics hold a fundamental importance in the development and fine-grained analysis of captioning systems. While current evaluation metrics tend to achieve an acceptable correlation with human judgeme…

Sentence

A Step Towards Mixture of Grader: Statistical Analysis of Existing Automatic Evaluation Metrics

2024-10-13 · Yun Joon Soh, Jishen Zhao

The explosion of open-sourced models and Question-Answering (QA) datasets emphasizes the importance of automated QA evaluation. We studied the statistics of the existing evaluation metrics for a better understanding of t…

Question Answering

NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist

2023-05-15 · Iftitahu Ni'mah, Meng Fang, Vlado Menkovski, Mykola Pechenizkiy

In this study, we analyze automatic evaluation metrics for Natural Language Generation (NLG), specifically task-agnostic metrics and human-aligned metrics. Task-agnostic metrics, such as Perplexity, BLEU, BERTScore, are …

Controllable Language ModellingDialogue GenerationLanguage Modellingnlg evaluation+3

Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge

2024-10-03 · Aparna Elangovan, Lei Xu, Jongwoo Ko, Mahsa Elyasi 외

The effectiveness of automatic evaluation of generative models is typically measured by comparing the labels generated via automation with labels by humans using correlation metrics. However, metrics like Krippendorff's …

Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

2022-04-21 · NAACL 2022 7 · Daniel Deutsch, Rotem Dror, Dan Roth

How reliably an automatic summarization evaluation metric replicates human judgments of summary quality is quantified by system-level correlations. We identify two ways in which the definition of the system-level correla…