paper-with-me

홈 › Papers

Pareto Frontiers in Deep Feature Learning: Data, Compute, Width, and Luck

2023-09-21 · NeurIPS 2023 11

In modern deep learning, algorithmic choices (such as width, depth, and learning rate) are known to modulate nuanced resource tradeoffs. This work investigates how these complexities necessarily arise for feature learning in the presence of computational-statistical gaps. We begin by considering offline sparse parity learning, a supervised classification problem which admits a statistical query lower bound for gradient-based training of a multilayer perceptron. This lower bound can be interpreted as a *multi-resource tradeoff frontier*: successful learning can only occur if one is sufficiently rich (large model), knowledgeable (large dataset), patient (many training iterations), or lucky (many random guesses). We show, theoretically and experimentally, that sparse initialization and increasing network width yield significant improvements in sample efficiency in this setting. Here, width plays the role of parallel search: it amplifies the probability of finding "lottery ticket" neurons, which learn sparse features more sample-efficiently. Finally, we show that the synthetic sparse parity task can be useful as a proxy for real problems requiring axis-aligned feature learning. We demonstrate improved sample efficiency on tabular classification benchmarks by using wide, sparsely-initialized MLP models; these networks sometimes outperform tuned random forests.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Inference economics of language models

2025-06-05 · Ege Erdil

We develop a theoretical model that addresses the economic trade-off between cost per token versus serial token generation speed when deploying LLMs for inference at scale. Our model takes into account arithmetic, memory…

Pareto Frontiers in Neural Feature Learning: Data, Compute, Width, and Luck

2023-09-07 · Benjamin L. Edelman, Surbhi Goel, Sham Kakade, Eran Malach 외

In modern deep learning, algorithmic choices (such as width, depth, and learning rate) are known to modulate nuanced resource tradeoffs. This work investigates how these complexities necessarily arise for feature learnin…

tabular-classification

Compute-Accuracy Pareto Frontiers for Open-Source Reasoning Large Language Models

2025-12-31 · Ákos Prucs, Nara Csutora, Mátyás Antal, Márk Marosi arxiv

Large Language Models (LLMs) are demonstrating rapid improvements on complex reasoning benchmarks, particularly when allowed to utilize intermediate reasoning steps before converging on a final solution. However, current…

The new hybrid COAW method for solving multi-objective problems

2016-01-06 · Zeinab Borhanifar, Elham Shadkam

In this article using Cuckoo Optimization Algorithm and simple additive weighting method the hybrid COAW algorithm is presented to solve multi-objective problems. Cuckoo algorithm is an efficient and structured method fo…

GaussianPSL: Soft partitioning for complex PSL problem

2025-09-22 · Phuong Mai Dinh, Van-Nam Huynh arxiv

Many practical applications of multi-objective optimization (MOO), including engineering design, autonomous systems, and machine learning, often yield complex Pareto frontiers (e.g., discontinuous, degenerate, or non-con…