paper-with-me

홈 › Papers

Reproduction and Replication: A Case Study with Automatic Essay Scoring

2020-05-01 · LREC 2020 5 · Eva Huber, {\c{C}}a{\u{g}}r{\i} {\c{C}}{\"o}ltekin

As in many experimental sciences, reproducibility of experiments has gained ever more attention in the NLP community. This paper presents our reproduction efforts of an earlier study of automatic essay scoring (AES) for determining the proficiency of second language learners in a multilingual setting. We present three sets of experiments with different objectives. First, as prescribed by the LREC 2020 REPROLANG shared task, we rerun the original AES system using the code published by the original authors on the same dataset. Second, we repeat the same experiments on the same data with a different implementation. And third, we test the original system on a different dataset and a different language. Most of our findings are in line with the findings of the original paper. Nevertheless, there are some discrepancies between our results and the results presented in the original paper. We report and discuss these differences in detail. We further go into some points related to confirmation of research findings through reproduction, including the choice of the dataset, reporting and accounting for variability, use of appropriate evaluation metrics, and making code and data available. We also discuss the varying uses and differences between the terms reproduction and replication, and we argue that reproduction, the confirmation of conclusions through independent experiments in varied settings is more valuable than exact replication of the published values.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bringing replication and reproduction together with generalisability in NLP: Three reproduction studies for Target Dependent Sentiment Analysis

2018-06-13 · COLING 2018 8 · Andrew Moore, Paul Rayson

Lack of repeatability and generalisability are two significant threats to continuing scientific development in Natural Language Processing. Language models and learning methods are so complex that scientific conference p…

Sentiment Analysis

Assessing the Reliability and Validity of Large Language Models for Automated Assessment of Student Essays in Higher Education

2025-08-04 · Andrea Gaggioli, Giuseppe Casaburi, Leonardo Ercolani, Francesco Collova' 외 arxiv

This study investigates the reliability and validity of five advanced Large Language Models (LLMs), Claude 3.5, DeepSeek v2, Gemini 2.5, GPT-4, and Mistral 24B, for automated essay scoring in a real world higher educatio…

Automated Essay Scoring

Reproduction and Replication of an Adversarial Stylometry Experiment

2022-08-15 · Haining Wang, Patrick Juola, Allen Riddell

Maintaining anonymity while communicating using natural language remains a challenge. Standard authorship attribution techniques that analyze candidate authors' writing styles achieve uncomfortably high accuracy even whe…

Authorship AttributionTranslation

Offspring from Reproduction Problems: What Replication Failure Teaches Us

2013-08-01 · ACL 2013 8 · Antske Fokkens, Marieke van Erp, Marten Postma, Ted Pedersen 외
Named Entity Recognition (NER)

Catching Idiomatic Expressions in EFL Essays

2018-06-01 · WS 2018 6 · Michael Flor, Beata Beigman Klebanov

This paper presents an exploratory study on large-scale detection of idiomatic expressions in essays written by non-native speakers of English. We describe a computational search procedure for automatic detection of idio…