paper-with-me

홈 › Papers

A global analysis of metrics used for measuring performance in natural language processing

2022-04-25 · nlppower (ACL) 2022 5 · Kathrin Blagec, Georg Dorffner, Milad Moradi, Simon Ott, Matthias Samwald

Measuring the performance of natural language processing models is challenging. Traditionally used metrics, such as BLEU and ROUGE, originally devised for machine translation and summarization, have been shown to suffer from low correlation with human judgment and a lack of transferability to other tasks and languages. In the past 15 years, a wide range of alternative metrics have been proposed. However, it is unclear to what extent this has had an impact on NLP benchmarking efforts. Here we provide the first large-scale cross-sectional analysis of metrics used for measuring performance in natural language processing. We curated, mapped and systematized more than 3500 machine learning model performance results from the open repository 'Papers with Code' to enable a global and comprehensive analysis. Our results suggest that the large majority of natural language processing metrics currently used have properties that may result in an inadequate reflection of a models' performance. Furthermore, we found that ambiguities and inconsistencies in the reporting of metrics may lead to difficulties in interpreting and comparing model performances, impairing transparency and reproducibility in NLP research.

📄 PDF Abstract BibTeX arXiv:2204.11574

Code (1)

OpenBioLink/ITO 공식 구현

Tasks

BenchmarkingMachine Translation

Similar Papers 제목 키워드 기반

Mind the Gaps: Measuring Visual Artifacts in Dimensionality Reduction

2025-11-18 · Jaume Ros, Alessio Arleo, Fernando Paulovich arxiv

Dimensionality Reduction (DR) techniques are commonly used for the visual exploration and analysis of high-dimensional data due to their ability to project datasets of high-dimensional points onto the 2D plane. However, …

Dimensionality Reduction

A critical analysis of metrics used for measuring progress in artificial intelligence

2020-08-06 · Kathrin Blagec, Georg Dorffner, Milad Moradi, Matthias Samwald

Comparing model performances on benchmark datasets is an integral part of measuring and driving progress in artificial intelligence. A model's performance on a benchmark dataset is commonly assessed based on a single or …

Benchmarking

Common Metrics to Benchmark Human-Machine Teams (HMT): A Review

2020-08-11 · Praveen Damacharla, Ahmad Y. Javaid, Jennie J. Gallimore, Vijay K. Devabhaktuni

A significant amount of work is invested in human-machine teaming (HMT) across multiple fields. Accurately and effectively measuring system performance of an HMT is crucial for moving the design of these systems forward.…

Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy

2025-03-25 · Athiya Deviyani, Fernando Diaz

Meta-evaluation of automatic evaluation metrics -- assessing evaluation metrics themselves -- is crucial for accurately benchmarking natural language processing systems and has implications for scientific inquiry, produc…

Benchmarkingspeech-recognitionSpeech Recognition

Metrics to Quantify Global Consistency in Synthetic Medical Images

2023-08-01 · Daniel Scholz, Benedikt Wiestler, Daniel Rueckert, Martin J. Menten

Image synthesis is increasingly being adopted in medical image processing, for example for data augmentation or inter-modality image translation. In these critical applications, the generated images must fulfill a high s…

Data AugmentationImage Generation