paper-with-me

홈 › Papers

A Simple and Fast Baseline for Tuning Large XGBoost Models

2021-11-12 · Sanyam Kapoor, Valerio Perrone

XGBoost, a scalable tree boosting algorithm, has proven effective for many prediction tasks of practical interest, especially using tabular datasets. Hyperparameter tuning can further improve the predictive performance, but unlike neural networks, full-batch training of many models on large datasets can be time consuming. Owing to the discovery that (i) there is a strong linear relation between dataset size & training time, (ii) XGBoost models satisfy the ranking hypothesis, and (iii) lower-fidelity models can discover promising hyperparameter configurations, we show that uniform subsampling makes for a simple yet fast baseline to speed up the tuning of large XGBoost models using multi-fidelity hyperparameter optimization with data subsets as the fidelity dimension. We demonstrate the effectiveness of this baseline on large-scale tabular datasets ranging from $15-70\mathrm{GB}$ in size.

📄 PDF Abstract BibTeX arXiv:2111.06924

Code (0)

등록된 구현이 없습니다.

Tasks

Hyperparameter Optimization

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Very fast Bayesian Additive Regression Trees on GPU

2024-10-30 · Giacomo Petrillo

Bayesian Additive Regression Trees (BART) is a nonparametric Bayesian regression technique based on an ensemble of decision trees. It is part of the toolbox of many statisticians. The overall statistical quality of the r…

CPUGPUregression

Faster Boosting with Smaller Memory

2019-01-25 · NeurIPS 2019 12 · Julaiti Alafate, Yoav Freund

State-of-the-art implementations of boosting, such as XGBoost and LightGBM, can process large training sets extremely fast. However, this performance requires that the memory size is sufficient to hold a 2-3 multiple of …

Exploring Microstructural Dynamics in Cryptocurrency Limit Order Books: Better Inputs Matter More Than Stacking Another Hidden Layer

2025-06-06 · Haochuan Wang

Cryptocurrency price dynamics are driven largely by microstructural supply demand imbalances in the limit order book (LOB), yet the highly noisy nature of LOB data complicates the signal extraction process. Prior researc…

Feature Engineering

Tabular Data: Deep Learning is Not All You Need

2021-06-06 · ICML Workshop AutoML 2021 7 · Ravid Shwartz-Ziv, Amitai Armon

A key element in solving real-life data science problems is selecting the types of models to use. Tree ensemble models (such as XGBoost) are usually recommended for classification and regression problems with tabular dat…

AllAutoMLDeep LearningGeneral Classification

Quantile Extreme Gradient Boosting for Uncertainty Quantification

2023-04-23 · Xiaozhe Yin, Masoud Fallah-Shorshani, Rob McConnell, Scott Fruin 외

As the availability, size and complexity of data have increased in recent years, machine learning (ML) techniques have become popular for modeling. Predictions resulting from applying ML models are often used for inferen…

Decision MakingPrediction Intervalsquantile regressionregression+1