paper-with-me

Papers

Estimating the Algorithmic Variance of Randomized Ensembles via the Bootstrap

2019-07-20 · Miles E. Lopes

Although the methods of bagging and random forests are some of the most widely used prediction methods, relatively little is known about their algorithmic convergence. In particular, there are not many theoretical guarantees for deciding when an ensemble is "large enough" --- so that its accuracy is close to that of an ideal infinite ensemble. Due to the fact that bagging and random forests are randomized algorithms, the choice of ensemble size is closely related to the notion of "algorithmic variance" (i.e. the variance of prediction error due only to the training algorithm). In the present work, we propose a bootstrap method to estimate this variance for bagging, random forests, and related methods in the context of classification. To be specific, suppose the training dataset is fixed, and let the random variable $Err_t$ denote the prediction error of a randomized ensemble of size $t$. Working under a "first-order model" for randomized ensembles, we prove that the centered law of $Err_t$ can be consistently approximated via the proposed method as $t\to\infty$. Meanwhile, the computational cost of the method is quite modest, by virtue of an extrapolation technique. As a consequence, the method offers a practical guideline for deciding when the algorithmic fluctuations of $Err_t$ are negligible.

📄 PDF Abstract BibTeX arXiv:1907.08742

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Estimating a sharp convergence bound for randomized ensembles

2013-03-04 · Miles E. Lopes

When randomized ensembles such as bagging or random forests are used for binary classification, the prediction error of the ensemble tends to decrease and stabilize as the number of classifiers increases. However, the pr…

Binary ClassificationDensity EstimationPrediction

Measuring the Algorithmic Convergence of Randomized Ensembles: The Regression Setting

2019-08-04 · Miles E. Lopes, Suofei Wu, Thomas C. M. Lee

When randomized ensemble methods such as bagging and random forests are implemented, a basic question arises: Is the ensemble large enough? In particular, the practitioner desires a rigorous guarantee that a given ensemb…

General ClassificationregressionVariable Selection

To Bag is to Prune

2020-08-17 · Philippe Goulet Coulombe

It is notoriously difficult to build a bad Random Forest (RF). Concurrently, RF blatantly overfits in-sample without any apparent consequence out-of-sample. Standard arguments, like the classic bias-variance trade-off or…

Bootstrap Inference for Quantile Treatment Effects in Randomized Experiments with Matched Pairs

2020-05-25 · Liang Jiang, Xiaobin Liu, Peter C. B. Phillips, Yichong Zhang

This paper examines methods of inference concerning quantile treatment effects (QTEs) in randomized experiments with matched-pairs designs (MPDs). Standard multiplier bootstrap inference fails to capture the negative dep…

Don't Explain Noise: Robust Counterfactuals for Randomized Ensembles

2022-05-27 · Alexandre Forel, Axel Parmentier, Thibaut Vidal

Counterfactual explanations describe how to modify a feature vector in order to flip the outcome of a trained classifier. Obtaining robust counterfactual explanations is essential to provide valid algorithmic recourse an…

counterfactualvalid