Closed-Form Beta Distribution Estimation from Sparse Statistics with Random Forest Implicit Regularization
This work advances distribution recovery from sparse data and ensemble classification through three main contributions. First, we introduce a closed-form estimator that reconstructs scaled beta distributions from limited statistics (minimum, maximum, mean, and median) via composite quantile and moment matching. The recovered parameters $(α,β)$, when used as features in Random Forest classifiers, improve pairwise classification on time-series snapshots, validating the fidelity of the recovered distributions. Second, we establish a link between classification accuracy and distributional closeness by deriving error bounds that constrain total variation distance and Jensen-Shannon divergence, the latter exhibiting quadratic convergence. Third, we show that zero-variance features act as an implicit regularizer, increasing selection probability for mid-ranked predictors and producing deeper, more varied trees. A SeatGeek pricing dataset serves as the primary application, illustrating distributional recovery and event-level classification while situating these methods within the structure and dynamics of the secondary ticket marketplace. The UCI handwritten digits dataset confirms the broader regularization effect. Overall, the study outlines a practical route from sparse distributional snapshots to closed-form estimation and improved ensemble accuracy, with reliability enhanced through implicit regularization.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Transformer-based Parameter Estimation in Statistics
Parameter estimation is one of the most important tasks in statistics, and is key to helping people understand the distribution behind a sample of observations. Traditionally parameter estimation is done either by closed…
Formparameter estimationMonte Carlo Simulation for Lasso-Type Problems by Estimator Augmentation
Regularized linear regression under the $\ell_1$ penalty, such as the Lasso, has been shown to be effective in variable selection and sparse modeling. The sampling distribution of an $\ell_1$-penalized estimator $\hat{\b…
regressionVariable SelectionVocal Bursts Type PredictionA Closed-Form Approximation to the Conjugate Prior of the Dirichlet and Beta Distributions
We derive the conjugate prior of the Dirichlet and beta distributions and explore it with numerical examples to gain an intuitive understanding of the distribution itself, its hyperparameters, and conditions concerning i…
FormSparse Estimation with Generalized Beta Mixture and the Horseshoe Prior
In this paper, the use of the Generalized Beta Mixture (GBM) and Horseshoe distributions as priors in the Bayesian Compressive Sensing framework is proposed. The distributions are considered in a two-layer hierarchical m…
Compressive SensingFast Capacity Estimation in Ultra-dense Wireless Networks with Random Interference
In wireless communication systems, the accurate and reliable evaluation of channel capacity is believed to be a fundamental and critical issue for terminals. However, with the rapid development of wireless technology, la…
Capacity Estimation