paper-with-me

홈 › Papers

A Reassessment of Reference-Based Grammatical Error Correction Metrics

2018-08-01 · COLING 2018 8 · Shamil Chollampatt, Hwee Tou Ng

Several metrics have been proposed for evaluating grammatical error correction (GEC) systems based on grammaticality, fluency, and adequacy of the output sentences. Previous studies of the correlation of these metrics with human quality judgments were inconclusive, due to the lack of appropriate significance tests, discrepancies in the methods, and choice of datasets used. In this paper, we re-evaluate reference-based GEC metrics by measuring the system-level correlations with humans on a large dataset of human judgments of GEC outputs, and by properly conducting statistical significance tests. Our results show no significant advantage of GLEU over MaxMatch (M2), contradicting previous studies that claim GLEU to be superior. For a finer-grained analysis, we additionally evaluate these metrics for their agreement with human judgments at the sentence level. Our sentence-level analysis indicates that comparing GLEU and M2, one metric may be more useful than the other depending on the scenario. We further qualitatively analyze these metrics and our findings show that apart from being less interpretable and non-deterministic, GLEU also produces counter-intuitive scores in commonly occurring test examples.

📄 PDF Abstract BibTeX

Code (1)

nusnlp/gecmetrics 공식 구현

Tasks

Grammatical Error CorrectionMachine TranslationSentence

Similar Papers 제목 키워드 기반

Reference-based Metrics can be Replaced with Reference-less Metrics in Evaluating Grammatical Error Correction Systems

2017-11-01 · IJCNLP 2017 11 · Hiroki Asano, Tomoya Mizumoto, Kentaro Inui

In grammatical error correction (GEC), automatically evaluating system outputs requires gold-standard references, which must be created manually and thus tend to be both expensive and limited in coverage. To address this…

Grammatical Error CorrectionMachine Translation

There's No Comparison: Reference-less Evaluation Metrics in Grammatical Error Correction

2016-10-07 · EMNLP 2016 11 · Courtney Napoles, Keisuke Sakaguchi, Joel Tetreault

Current methods for automatically evaluating grammatical error correction (GEC) systems rely on gold-standard references. However, these methods suffer from penalizing grammatical edits that are correct but not in the go…

BenchmarkingGrammatical Error CorrectionSentence

CLEME2.0: Towards More Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction

2024-07-01 · Jingheng Ye, Zishan Xu, Yinghui Li, Xuxin Cheng 외

The paper focuses on improving the interpretability of Grammatical Error Correction (GEC) metrics, which receives little attention in previous studies. To bridge the gap, we propose CLEME2.0, a reference-based evaluation…

Grammatical Error Correction

Reliability Crisis of Reference-free Metrics for Grammatical Error Correction

2025-09-30 · Takumi Goto, Yusuke Sakai, Taro Watanabe arxiv

Reference-free evaluation metrics for grammatical error correction (GEC) have achieved high correlation with human judgments. However, these metrics are not designed to evaluate adversarial systems that aim to obtain unj…

Grammatical Error CorrectionAdversarial Attack

SOME: Reference-less Sub-Metrics Optimized for Manual Evaluations of Grammatical Error Correction

2020-12-01 · COLING 2020 8 · Ryoma Yoshimura, Masahiro Kaneko, Tomoyuki Kajiwara, Mamoru Komachi

We propose a reference-less metric trained on manual evaluations of system outputs for grammatical error correction (GEC). Previous studies have shown that reference-less metrics are promising; however, existing metrics …

Grammatical Error CorrectionSentence