paper-with-me

홈 › Papers

The Prevalence of Errors in Machine Learning Experiments

2019-09-10 · Martin Shepperd, Yuchen Guo, Ning li, Mahir Arzoky, Andrea Capiluppi, Steve Counsell, Giuseppe Destefanis, Stephen Swift, Allan Tucker, Leila Yousefi

Context: Conducting experiments is central to research machine learning research to benchmark, evaluate and compare learning algorithms. Consequently it is important we conduct reliable, trustworthy experiments. Objective: We investigate the incidence of errors in a sample of machine learning experiments in the domain of software defect prediction. Our focus is simple arithmetical and statistical errors. Method: We analyse 49 papers describing 2456 individual experimental results from a previously undertaken systematic review comparing supervised and unsupervised defect prediction classifiers. We extract the confusion matrices and test for relevant constraints, e.g., the marginal probabilities must sum to one. We also check for multiple statistical significance testing errors. Results: We find that a total of 22 out of 49 papers contain demonstrable errors. Of these 7 were statistical and 16 related to confusion matrix inconsistency (one paper contained both classes of error). Conclusions: Whilst some errors may be of a relatively trivial nature, e.g., transcription errors their presence does not engender confidence. We strongly urge researchers to follow open science principles so errors can be more easily be detected and corrected, thus as a community reduce this worryingly high error rate with our computational experiments.

📄 PDF Abstract BibTeX arXiv:1909.04436

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine Learning

Similar Papers 제목 키워드 기반

Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

2021-03-26 · Curtis G. Northcutt, Anish Athalye, Jonas Mueller

We identify label errors in the test sets of 10 of the most commonly-used computer vision, natural language, and audio datasets, and subsequently study the potential for these label errors to affect benchmark results. Er…

BIG-bench Machine Learning

Association Between Neighborhood Factors and Adult Obesity in Shelby County, Tennessee: Geospatial Machine Learning Approach

2022-08-09 · Whitney S Brakefield, Olufunto A Olusanya, Arash Shaban-Nejad

Obesity is a global epidemic causing at least 2.8 million deaths per year. This complex disease is associated with significant socioeconomic burden, reduced work productivity, unemployment, and other social determinants …

Decision Making

Quantifying disparities in intimate partner violence: a machine learning method to correct for underreporting

2021-10-08 · Divya Shanmugam, Kaihua Hou, Emma Pierson

Estimating the prevalence of a medical condition, or the proportion of the population in which it occurs, is a fundamental problem in healthcare and public health. Accurate estimates of the relative prevalence across gro…

Diagnosis Prevalence vs. Efficacy in Machine-learning Based Diagnostic Decision Support

2020-06-24 · Gil Alon, Elizabeth Chen, Guergana Savova, Carsten Eickhoff

Many recent studies use machine learning to predict a small number of ICD-9-CM codes. In practice, on the other hand, physicians have to consider a broader range of diagnoses. This study aims to put these previously inco…

BIG-bench Machine LearningDiagnostic

Deployment of Image Analysis Algorithms under Prevalence Shifts

2023-03-22 · Patrick Godau, Piotr Kalinowski, Evangelia Christodoulou, Annika Reinke 외

Domain gaps are among the most relevant roadblocks in the clinical translation of machine learning (ML)-based solutions for medical image analysis. While current research focuses on new training paradigms and network arc…

image-classificationImage ClassificationMedical Image Analysis