paper-with-me

Papers

Reproducibility in NLP: What Have We Learned from the Checklist?

2023-06-16 · Ian Magnusson, Noah A. Smith, Jesse Dodge

Scientific progress in NLP rests on the reproducibility of researchers' claims. The *CL conferences created the NLP Reproducibility Checklist in 2020 to be completed by authors at submission to remind them of key information to include. We provide the first analysis of the Checklist by examining 10,405 anonymous responses to it. First, we find evidence of an increase in reporting of information on efficiency, validation performance, summary statistics, and hyperparameters after the Checklist's introduction. Further, we show acceptance rate grows for submissions with more Yes responses. We find that the 44% of submissions that gather new data are 5% less likely to be accepted than those that did not; the average reviewer-rated reproducibility of these submissions is also 2% lower relative to the rest. We find that only 46% of submissions claim to open-source their code, though submissions that do have 8% higher reproducibility score relative to those that do not, the most for any item. We discuss what can be inferred about the state of reproducibility in NLP, and provide a set of recommendations for future conferences, including: a) allowing submitting code and appendices one week after the deadline, and b) measuring dataset reproducibility by a checklist of data collection practices.

📄 PDF Abstract BibTeX arXiv:2306.09562

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Shift Toward Open and Reproducible AI Research

2026-06-15 · Kevin L Coakley, Thijs Snelleman, Holger Hoos, Odd Erik Gundersen arxiv

The reproducibility crisis has directed the AI research community toward improving documentation practices. Several studies have identified methodological issues, and in response, the most impactful venues in the field h…

Long-term Reproducibility for Neural Architecture Search

2022-07-11 · David Towers, Matthew Forshaw, Amir Atapour-Abarghouei, Andrew Stephen McGough

It is a sad reflection of modern academia that code is often ignored after publication -- there is no academic 'kudos' for bug fixes / maintenance. Code is often unavailable or, if available, contains bugs, is incomplete…

Neural Architecture Search

Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program)

2020-03-27 · Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vincent Larivière 외

One of the challenges in machine learning research is to ensure that presented and published results are sound and reliable. Reproducibility, that is obtaining similar results as presented in a paper or talk, using the s…

BIG-bench Machine Learning

RIDGE: Reproducibility, Integrity, Dependability, Generalizability, and Efficiency Assessment of Medical Image Segmentation Models

2024-01-16 · Farhad Maleki, Linda Moy, Reza Forghani, Tapotosh Ghosh 외

Deep learning techniques hold immense promise for advancing medical image analysis, particularly in tasks like image segmentation, where precise annotation of regions or volumes of interest within medical images is cruci…

Deep LearningImage SegmentationMedical Image AnalysisMedical Image Segmentation+3

Assessing Reproducibility in Evolutionary Computation: A Case Study using Human- and LLM-based Assessment

2026-02-05 · Francesca Da Ros, Tarik Začiragić, Aske Plaat, Thomas Bäck 외 arxiv

Reproducibility is an important requirement in evolutionary computation, where results largely depend on computational experiments. In practice, reproducibility relies on how algorithms, experimental protocols, and artif…