Horseshoe Regularization for Feature Subset Selection
Feature subset selection arises in many high-dimensional applications of statistics, such as compressed sensing and genomics. The $\ell_0$ penalty is ideal for this task, the caveat being it requires the NP-hard combinatorial evaluation of all models. A recent area of considerable interest is to develop efficient algorithms to fit models with a non-convex $\ell_\gamma$ penalty for $\gamma\in (0,1)$, which results in sparser models than the convex $\ell_1$ or lasso penalty, but is harder to fit. We propose an alternative, termed the horseshoe regularization penalty for feature subset selection, and demonstrate its theoretical and computational advantages. The distinguishing feature from existing non-convex optimization approaches is a full probabilistic representation of the penalty as the negative of the logarithm of a suitable prior, which in turn enables efficient expectation-maximization and local linear approximation algorithms for optimization and MCMC for uncertainty quantification. In synthetic and real data, the resulting algorithms provide better statistical performance, and the computation requires a fraction of time of state-of-the-art non-convex solvers.
Code (1)
Tasks
compressed sensingUncertainty QuantificationSimilar Papers 제목 키워드 기반
Horseshoe Regularization for Machine Learning in Complex and Deep Models
Since the advent of the horseshoe priors for regularization, global-local shrinkage methods have proved to be a fertile ground for the development of Bayesian methodology in machine learning, specifically for high-dimens…
BIG-bench Machine LearningregressionFlexible metro network architecture based on wavelength blockers and coherent transmission
A flexible semi-filterless architecture for metro DWDM networks with ring- or horseshoe topologies is presented. ROADM nodes are based on wavelength blockers. Coherent transmission enables filterless channel selection. C…
channel selectionHorseshoe Mixtures-of-Experts (HS-MoE)
Horseshoe mixtures-of-experts (HS-MoE) models provide a Bayesian framework for sparse expert selection in mixture-of-experts architectures. We combine the horseshoe prior's adaptive global-local shrinkage with input-depe…
Model Selection in Bayesian Neural Networks via Horseshoe Priors
Bayesian Neural Networks (BNNs) have recently received increasing attention for their ability to provide well-calibrated posterior uncertainties. However, model selection---even choosing the number of nodes---remains an …
Model SelectionOpen-Ended Question AnsweringFeature selection via simultaneous sparse approximation for person specific face verification
There is an increasing use of some imperceivable and redundant local features for face recognition. While only a relatively small fraction of them is relevant to the final recognition task, the feature selection is a cru…
Face RecognitionFace Verificationfeature selection