paper-with-me

Papers

The ReproGen Shared Task on Reproducibility of Human Evaluations in NLG: Overview and Results

2021-08-01 · INLG (ACL) 2021 8 · Anya Belz, Anastasia Shimorina, Shubham Agarwal, Ehud Reiter

The NLP field has recently seen a substantial increase in work related to reproducibility of results, and more generally in recognition of the importance of having shared definitions and practices relating to evaluation. Much of the work on reproducibility has so far focused on metric scores, with reproducibility of human evaluation results receiving far less attention. As part of a research programme designed to develop theory and practice of reproducibility assessment in NLP, we organised the first shared task on reproducibility of human evaluations, ReproGen 2021. This paper describes the shared task in detail, summarises results from each of the reproduction studies submitted, and provides further comparative analysis of the results. Out of nine initial team registrations, we received submissions from four teams. Meta-analysis of the four reproduction studies revealed varying degrees of reproducibility, and allowed very tentative first conclusions about what types of evaluation tend to have better reproducibility.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ReproGen: Proposal for a Shared Task on Reproducibility of Human Evaluations in NLG

2020-12-01 · INLG (ACL) 2020 12 · Anya Belz, Shubham Agarwal, Anastasia Shimorina, Ehud Reiter

Across NLP, a growing body of work is looking at the issue of reproducibility. However, replicability of human evaluation experiments and reproducibility of their results is currently under-addressed, and this is of part…

TUDA-Reproducibility @ ReproGen: Replicability of Human Evaluation of Text-to-Text and Concept-to-Text Generation

2021-08-01 · INLG (ACL) 2021 8 · Christian Richter, Yanran Chen, Steffen Eger

This paper describes our contribution to the Shared Task ReproGen by Belz et al. (2021), which investigates the reproducibility of human evaluations in the context of Natural Language Generation. We selected the paper “G…

Concept-To-Text GenerationPaper generationText Generation

Another PASS: A Reproduction Study of the Human Evaluation of a Football Report Generation System

2021-08-01 · INLG (ACL) 2021 8 · Simon Mille, Thiago castro Ferreira, Anya Belz, Brian Davis

This paper reports results from a reproduction study in which we repeated the human evaluation of the PASS Dutch-language football report generation system (van der Lee et al., 2017). The work was carried out as part of …

A Reproduction Study of an Annotation-based Human Evaluation of MT Outputs

2021-08-01 · INLG (ACL) 2021 8 · Maja Popović, Anya Belz

In this paper we report our reproduction study of the Croatian part of an annotation-based human evaluation of machine-translated user reviews (Popovic, 2020). The work was carried out as part of the ReproGen Shared Task…

Experimental Design

Reproducing a Comparison of Hedged and Non-hedged NLG Texts

2021-08-01 · INLG (ACL) 2021 8 · Saad Mahamood

This paper describes an attempt to reproduce an earlier experiment, previously conducted by the author, that compares hedged and non-hedged NLG texts as part of the ReproGen shared challenge. This reproduction effort was…