paper-with-me

홈 › Papers

A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection

2026-06-29 · Andrey A. Dukhovny, Andrey M. Lange arxiv

The number of trees is a central computational parameter in Random Forests: increasing it reduces finite-ensemble variability but increases training and prediction cost. Plateau-based tuning adapts this parameter through local comparisons of out-of-bag scores at a geometric triplet of tree counts. After the remaining hyperparameters have stabilized, however, the central triplet point need not converge to a deterministic value; instead, it fluctuates around a stationary regime. This paper develops a stationary-distribution theory for this process. The central ensemble size $B_t$ is modeled as a birth-death Markov chain on a geometric grid, and its stationary distribution is derived through local balance. Under a leading centered folded-normal approximation, equilibrium equations are obtained for the original update rule and a symmetric modified variant, implying that the stationary center $B_*=O(\varepsilon^{-2})$ as $\varepsilon\downarrow 0$. The stationary spread is also characterized. A local Gaussian approximation and a Fokker-Planck interpretation give grid-level variance constants. After conversion to the ensemble-size scale, $σ_{B,*}=O(\varepsilon^{-2})$, while the variance is $O(\varepsilon^{-4})$. The leading relative spread is independent of $\varepsilon$ and controlled by the scale factor and update rule. These results interpret plateau-based Random Forest tuning as a stochastic process rather than a deterministic stopping rule.

📄 PDF Abstract BibTeX arXiv:2606.30837

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Many Trees in a Random Forest? A Revisited Approach with Plateau Search and Optuna Integration

2026-06-02 · Vadim Porvatov, Andrey Dukhovny, Andrey Lange arxiv

Hyperparameter optimization (HPO) for Random Forest faces a specific difficulty in tuning the number of trees: the predictive score typically improves monotonically with ensemble size, so standard methods such as Tree-st…

Hyperparameter Optimization

A Geometric Characterization of the Stationary Plateau for Two-Layer Neural Networks

2026-06-03 · Tian Ding, Dawei Li, Ruoyu Sun arxiv

We investigate the geometric structure of stationary plateaus that arise in the loss landscape of two-layer neural networks with smooth activation functions. We focus on the phenomenon of "neuron splitting" where duplica…

Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis

2020-01-10 · NeurIPS 2019 12 · Yuki Yoshida, Masato Okada

The plateau phenomenon, wherein the loss value stops decreasing during the process of learning, has been reported by various researchers. The phenomenon is actively inspected in the 1990s and found to be due to the funda…

TripletGAN: Training Generative Model with Triplet Loss

2017-11-14 · Gongze Cao, Yezhou Yang, Jie Lei, Cheng Jin 외

As an effective way of metric learning, triplet loss has been widely used in many deep learning tasks, including face recognition and person-ReID, leading to many states of the arts. The main innovation of triplet loss i…

Face RecognitionGeneral ClassificationMetric Learningmodel+1

Huber Additive Models for Non-stationary Time Series Analysis

2021-09-29 · ICLR 2022 4 · Yingjie Wang, Xianrui Zhong, Fengxiang He, Hong Chen 외

Sparse additive models have shown promising flexibility and interpretability in processing time series data. However, existing methods usually assume the time series data to be stationary and the innovation is sampled fr…

Additive modelsCausal DiscoveryGeneralization BoundsLearning Theory+2