paper-with-me

홈 › Papers

The Backbone Method for Ultra-High Dimensional Sparse Machine Learning

2020-06-11 · Dimitris Bertsimas, Vassilis Digalakis Jr

We present the backbone method, a generic framework that enables sparse and interpretable supervised machine learning methods to scale to ultra-high dimensional problems. We solve sparse regression problems with $10^7$ features in minutes and $10^8$ features in hours, as well as decision tree problems with $10^5$ features in minutes.The proposed method operates in two phases: we first determine the backbone set, consisting of potentially relevant features, by solving a number of tractable subproblems; then, we solve a reduced problem, considering only the backbone features. For the sparse regression problem, our theoretical analysis shows that, under certain assumptions and with high probability, the backbone set consists of the truly relevant features. Numerical experiments on both synthetic and real-world datasets demonstrate that our method outperforms or competes with state-of-the-art methods in ultra-high dimensional problems, and competes with optimal solutions in problems where exact methods scale, both in terms of recovering the truly relevant features and in its out-of-sample predictive performance.

📄 PDF Abstract BibTeX arXiv:2006.06592

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningregressionVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

CSRv2: Unlocking Ultra-Sparse Embeddings

2026-02-05 · Lixuan Guo, Yifei Wang, Tiansheng Wen, Yifan Wang 외 arxiv

In the era of large foundation models, the quality of embeddings has become a central determinant of downstream task performance and overall system capability. Yet widely used dense embeddings are often extremely high-di…

Representation Learning

Large-scale Online Feature Selection for Ultra-high Dimensional Sparse Data

2014-09-27 · Yue Wu, Steven C. H. Hoi, Tao Mei, Nenghai Yu

Feature selection with large-scale high-dimensional data is important yet very challenging in machine learning and data mining. Online feature selection is a promising new paradigm that is more efficient and scalable tha…

feature selectionVocal Bursts Intensity Prediction

Graph Learning via Spectral Densification

2021-01-01 · Zhuo Feng, Yongyu Wang, Zhiqiang Zhao

Graph learning plays important role in many data mining and machine learning tasks, such as manifold learning, data representation and analysis, dimensionality reduction, data clustering, and visualization, etc. For the …

BIG-bench Machine LearningClusteringDimensionality ReductionGraph Learning

Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval

2021-04-15 · EMNLP 2021 11 · Kyoung-Rok Jang, Junmo Kang, Giwon Hong, Sung-Hyon Myaeng 외

The semantic matching capabilities of neural information retrieval can ameliorate synonymy and polysemy problems of symbolic approaches. However, neural models' dense representations are more suitable for re-ranking, due…

BinarizationInformation RetrievalLanguage ModellingRe-Ranking+3

BackboneLearn: A Library for Scaling Mixed-Integer Optimization-Based Machine Learning

2023-11-22 · Vassilis Digalakis Jr, Christos Ziakas

We present BackboneLearn: an open-source software package and framework for scaling mixed-integer optimization (MIO) problems with indicator variables to high-dimensional problems. This optimization paradigm can naturall…

Clustering