paper-with-me

Papers

Cross-validation failure: small sample sizes lead to large error bars

2017-06-23 · Gaël Varoquaux

Predictive models ground many state-of-the-art developments in statistical brain image analysis: decoding, MVPA, searchlight, or extraction of biomarkers. The principled approach to establish their validity and usefulness is cross-validation, testing prediction on unseen data. Here, I would like to raise awareness on error bars of cross-validation, which are often underestimated. Simple experiments show that sample sizes of many neuroimaging studies inherently lead to large error bars, eg $\pm$10% for 100 samples. The standard error across folds strongly underestimates them. These large error bars compromise the reliability of conclusions drawn with predictive models, such as biomarkers or methods developments where, unlike with cognitive neuroimaging MVPA approaches, more samples cannot be acquired by repeating the experiment across many subjects. Solutions to increase sample size must be investigated, tackling possible increases in heterogeneity of the data.

📄 PDF Abstract BibTeX arXiv:1706.07581

Code (1)

GaelVaroquaux/cross_validation_failure 공식 구현

Similar Papers 제목 키워드 기반

Extrapolated cross-validation for randomized ensembles

2023-02-27 · Jin-Hong Du, Pratik Patil, Kathryn Roeder, Arun Kumar Kuchibhotla

Ensemble methods such as bagging and random forests are ubiquitous in various fields, from finance to genomics. Despite their prevalence, the question of the efficient tuning of ensemble parameters has received relativel…

Evaluating Synthetic Tabular Data Generated To Augment Small Sample Datasets

2022-11-19 · Javier Marin

This work proposes a method to evaluate synthetic tabular data generated to augment small sample datasets. While data augmentation techniques can increase sample counts for machine learning applications, traditional vali…

Data AugmentationTopological Data Analysis

Combined Pruning for Nested Cross-Validation to Accelerate Automated Hyperparameter Optimization for Embedded Feature Selection in High-Dimensional Data with Very Small Sample Sizes

2022-02-01 · Sigrun May, Sven Hartmann, Frank Klawonn

Background: Embedded feature selection in high-dimensional data with very small sample sizes requires optimized hyperparameters for the model building process. For this hyperparameter optimization, nested cross-validatio…

feature selectionHyperparameter Optimization

Bayesian Safety Validation for Failure Probability Estimation of Black-Box Systems

2023-05-03 · Robert J. Moss, Mykel J. Kochenderfer, Maxime Gariel, Arthur Dubois

Estimating the probability of failure is an important step in the certification of safety-critical systems. Efficient estimation methods are often needed due to the challenges posed by high-dimensional input spaces, risk…

Bayesian OptimizationDecision Making

Statistical Agnostic Mapping: a Framework in Neuroimaging based on Concentration Inequalities

2019-12-27 · J M Gorriz, SiPBA Group, CAM neuroscience

In the 70s a novel branch of statistics emerged focusing its effort in selecting a function in the pattern recognition problem, which fulfils a definite relationship between the quality of the approximation and its compl…

Two-sample testing