Missing Information, Unresponsive Authors, Experimental Flaws: The Impossibility of Assessing the Reproducibility of Previous Human Evaluations in NLP
We report our efforts in identifying a set of previous human evaluations in NLP that would be suitable for a coordinated study examining what makes human evaluations in NLP more/less reproducible. We present our results and findings, which include that just 13\% of papers had (i) sufficiently low barriers to reproduction, and (ii) enough obtainable information, to be considered for reproduction, and that all but one of the experiments we selected for reproduction was discovered to have flaws that made the meaningfulness of conducting a reproduction questionable. As a result, we had to change our coordinated study design from a reproduce approach to a standardise-then-reproduce-twice approach. Our overall (negative) finding that the great majority of human evaluations in NLP is not repeatable and/or not reproducible and/or too flawed to justify reproduction, paints a dire picture, but presents an opportunity for a rethink about how to design and report human evaluations in NLP.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Comments on Mathematical Modeling of Current Source Matrix Converter with Venturini and SVM
In this paper, authors want to comment on a recently published article describing the Mathematical Modeling of Current Source Matrix Converter (CSMC) with two modulation strategies, namely: Venturini and Space Vector Mod…
Perfectly predicting ICU length of stay: too good to be true
A paper of Alsinglawi et al was recently accepted and published in Scientific Reports. In this paper, the authors aim to predict length of stay (LOS), discretized into either long (> 7 days) or short stays (< 7 days), of…
ManagementAre We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
Scholarly peer review is a cornerstone of scientific advancement, but the system is under strain due to increasing manuscript submissions and the labor-intensive nature of the process. Recent advancements in large langua…
On the Importance of Strong Baselines in Bayesian Deep Learning
Like all sub-fields of machine learning Bayesian Deep Learning is driven by empirical validation of its theoretical proposals. Given the many aspects of an experiment it is always possible that minor or even major experi…
Deep LearningWhat's under the hood: Investigating Automatic Metrics on Meeting Summarization
Meeting summarization has become a critical task considering the increase in online interactions. While new techniques are introduced regularly, their evaluation uses metrics not designed to capture meeting-specific erro…
Meeting Summarization