paper-with-me

홈 › Papers

Decision trees as partitioning machines to characterize their generalization properties

2020-10-14 · NeurIPS 2020 12 · Jean-Samuel Leboeuf, Frédéric LeBlanc, Mario Marchand

Decision trees are popular machine learning models that are simple to build and easy to interpret. Even though algorithms to learn decision trees date back to almost 50 years, key properties affecting their generalization error are still weakly bounded. Hence, we revisit binary decision trees on real-valued features from the perspective of partitions of the data. We introduce the notion of partitioning function, and we relate it to the growth function and to the VC dimension. Using this new concept, we are able to find the exact VC dimension of decision stumps, which is given by the largest integer $d$ such that $2\ell \ge \binom{d}{\left\lfloor\frac{d}{2}\right\rfloor}$, where $\ell$ is the number of real-valued features. We provide a recursive expression to bound the partitioning functions, resulting in a upper bound on the growth function of any decision tree structure. This allows us to show that the VC dimension of a binary tree structure with $N$ internal nodes is of order $N \log(N\ell)$. Finally, we elaborate a pruning algorithm based on these results that performs better than the CART algorithm on a number of datasets, with the advantage that no cross-validation is required.

📄 PDF Abstract BibTeX arXiv:2010.07374

Code (1)

jsleb333/paper-decision-trees-as-partitioning-machines 공식 구현

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Feature Learning for Interpretable, Performant Decision Trees

2023-09-21 · NeurIPS 2023 11

Decision trees are regarded for high interpretability arising from their hierarchical partitioning structure built on simple decision rules. However, in practice, this is not realized because axis-aligned partitioning of…

Decision Machines: Congruent Decision Trees

2021-01-27 · Jinxiong Zhang

The decision tree recursively partitions the input space into regions and derives axis-aligned decision boundaries from data. Despite its simplicity and interpretability, decision trees lack parameterized representation,…

Computational Efficiency

The return of AdaBoost.MH: multi-class Hamming trees

2013-12-20 · Balázs Kégl

Within the framework of AdaBoost.MH, we propose to train vector-valued decision trees to optimize the multi-class edge without reducing the multi-class problem to $K$ binary one-against-all classifications. The key eleme…

Computational Efficiency

Guided Random Forest and its application to data approximation

2019-09-02 · Prashant Gupta, Aashi Jindal, Jayadeva, Debarka Sengupta

We present a new way of constructing an ensemble classifier, named the Guided Random Forest (GRAF) in the sequel. GRAF extends the idea of building oblique decision trees with localized partitioning to obtain a global pa…

Yggdrasil: An Optimized System for Training Deep Decision Trees at Scale

2016-12-01 · NeurIPS 2016 12 · Firas Abuzaid, Joseph K. Bradley, Feynman T. Liang, Andrew Feng 외

Deep distributed decision trees and tree ensembles have grown in importance due to the need to model increasingly large datasets. However, PLANET, the standard distributed tree learning algorithm implemented in systems …

CPU