paper-with-me

홈 › Papers

There is no Double-Descent in Random Forests

2021-11-08 · Sebastian Buschjäger, Katharina Morik

Random Forests (RFs) are among the state-of-the-art in machine learning and offer excellent performance with nearly zero parameter tuning. Remarkably, RFs seem to be impervious to overfitting even though their basic building blocks are well-known to overfit. Recently, a broadly received study argued that a RF exhibits a so-called double-descent curve: First, the model overfits the data in a u-shaped curve and then, once a certain model complexity is reached, it suddenly improves its performance again. In this paper, we challenge the notion that model capacity is the correct tool to explain the success of RF and argue that the algorithm which trains the model plays a more important role than previously thought. We show that a RF does not exhibit a double-descent curve but rather has a single descent. Hence, it does not overfit in the classic sense. We further present a RF variation that also does not overfit although its decision boundary approximates that of an overfitted DT. Similar, we show that a DT which approximates the decision boundary of a RF will still overfit. Last, we study the diversity of an ensemble as a tool the estimate its performance. To do so, we introduce Negative Correlation Forest (NCForest) which allows for precise control over the diversity in the ensemble. We show, that the diversity and the bias indeed have a crucial impact on the performance of the RF. Having too low diversity collapses the performance of the RF into a a single tree, whereas having too much diversity means that most trees do not produce correct outputs anymore. However, in-between these two extremes we find a large range of different trade-offs with all roughly equal performance. Hence, the specific trade-off between bias and diversity does not matter as long as the algorithm reaches this good trade-off regime.

📄 PDF Abstract BibTeX arXiv:2111.04409

Code (1)

sbuschjaeger/rf-double-descent 공식 구현 pytorch

Tasks

Diversity

Similar Papers 제목 키워드 기반

Trees, Forests, Chickens, and Eggs: When and Why to Prune Trees in a Random Forest

2021-03-30 · Siyu Zhou, Lucas Mentch

Due to their long-standing reputation as excellent off-the-shelf predictors, random forests continue remain a go-to model of choice for applied statisticians and data scientists. Despite their widespread use, however, un…

Random Hinge Forest for Differentiable Learning

2018-02-12 · Nathan Lay, Adam P. Harrison, Sharon Schreiber, Gitesh Dawer 외

We propose random hinge forests, a simple, efficient, and novel variant of decision forests. Importantly, random hinge forests can be readily incorporated as a general component within arbitrary computation graphs that a…

A Rigorous, Tractable Measure of Model Complexity

2026-05-20 · Oskar Allerbo, Thomas B. Schön arxiv

An accurate assessment of a model's complexity is crucial for topics such as interpretation, generalization, and model selection. However, most existing complexity measures either rely on heuristic assumptions or are com…

Oblique and rotation double random forest

2021-11-03 · M. A. Ganaie, M. Tanveer, P. N. Suganthan, V. Snasel

Random Forest is an ensemble of decision trees based on the bagging and random subspace concepts. As suggested by Breiman, the strength of unstable learners and the diversity among them are the ensemble models' core stre…

Diversity

High-dimensional analysis of double descent for linear regression with random projections

2023-03-02 · Francis Bach

We consider linear regression problems with a varying number of random projections, where we provably exhibit a double descent curve for a fixed prediction problem, with a high-dimensional analysis based on random matrix…

regression