paper-with-me

홈 › Papers

The Unreasonable Effectiveness of Open Science in AI: A Replication Study

2024-12-20 · Odd Erik Gundersen, Odd Cappelen, Martin Mølnå, Nicklas Grimstad Nilsen

A reproducibility crisis has been reported in science, but the extent to which it affects AI research is not yet fully understood. Therefore, we performed a systematic replication study including 30 highly cited AI studies relying on original materials when available. In the end, eight articles were rejected because they required access to data or hardware that was practically impossible to acquire as part of the project. Six articles were successfully reproduced, while five were partially reproduced. In total, 50% of the articles included was reproduced to some extent. The availability of code and data correlate strongly with reproducibility, as 86% of articles that shared code and data were fully or partly reproduced, while this was true for 33% of articles that shared only data. The quality of the data documentation correlates with successful replication. Poorly documented or miss-specified data will probably result in unsuccessful replication. Surprisingly, the quality of the code documentation does not correlate with successful replication. Whether the code is poorly documented, partially missing, or not versioned is not important for successful replication, as long as the code is shared. This study emphasizes the effectiveness of open science and the importance of properly documenting data work.

📄 PDF Abstract BibTeX arXiv:2412.17859

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

Gricea: An Open Science Platform for Conversational AI Research

2026-09-18 · Nikhil Sharma, Yunlin Gong, Xinyang Cheng, Ziang Xiao hf

We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge accum…

Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater Reliability

2021-08-01 · ACL 2021 5 · Ka Wong, Praveen Paritosh, Lora Aroyo

When collecting annotations and labeled data from humans, a standard practice is to use inter-rater reliability (IRR) as a measure of data goodness (Hallgren, 2012). Metrics such as Krippendorff{'}s alpha or Cohen{'}s ka…

Benchmarking

In-class Data Analysis Replications: Teaching Students while Testing Science

2023-08-31 · Kristina Gligoric, Tiziano Piccardi, Jake Hofman, Robert West

Science is facing a reproducibility crisis. Previous work has proposed incorporating data analysis replications into classrooms as a potential solution. However, despite the potential benefits, it is unclear whether this…

Replication Markets: Results, Lessons, Challenges and Opportunities in AI Replication

2020-05-10 · Yang Liu, Michael Gordon, Juntao Wang, Michael Bishop 외

The last decade saw the emergence of systematic large-scale replication projects in the social and behavioral sciences, (Camerer et al., 2016, 2018; Ebersole et al., 2016; Klein et al., 2014, 2018; Collaboration, 2015). …

The positive-negative mode link between brain connectivity, demographics, and behavior: A pre-registered replication of Smith et al. 2015

2022-01-25 · Nikhil Goyal1, Dustin Moraczewski, Peter A. Bandettini, Emily S. Finn 외

In mental health research, it has proven difficult to find measures of brain function that provide reliable indicators of mental health and well-being, including susceptibility to mental health disorders. Recently, a fam…

Functional Connectivity