paper-with-me

Papers

Optimal scaling laws in learning hierarchical multi-index models

2026-02-05 · Leonardo Defilippis, Florent Krzakala, Bruno Loureiro, Antoine Maillard arxiv

In this work, we provide a sharp theory of scaling laws for two-layer neural networks trained on a class of hierarchical multi-index targets, in a genuinely representation-limited regime. We derive exact information-theoretic scaling laws for subspace recovery and prediction error, revealing how the hierarchical features of the target are sequentially learned through a cascade of phase transitions. We further show that these optimal rates are achieved by a simple, target-agnostic spectral estimator, which can be interpreted as the small learning-rate limit of gradient descent on the first-layer weights. Once an adapted representation is identified, the readout can be learned statistically optimally, using an efficient procedure. As a consequence, we provide a unified and rigorous explanation of scaling laws, plateau phenomena, and spectral structure in shallow neural networks trained on such hierarchical targets.

📄 PDF Abstract BibTeX arXiv:2602.05846

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design

2026-02-10 · Bojian Hou, Xiaolong Liu, Xiaoyi Liu, Jiaqi Xu 외 arxiv

Deriving predictable scaling laws that govern the relationship between model performance and computational investment is crucial for designing and allocating resources in massive-scale recommendation systems. While such …

Recommendation Systems

Generalizing Scaling Laws for Dense and Sparse Large Language Models

2025-08-08 · Md Arafat Hossain, Xingfu Wu, Valerie Taylor, Ali Jannesari arxiv

Despite recent advancements of large language models (LLMs), optimally predicting the model size for LLM pretraining or allocating optimal resources still remains a challenge. Several efforts have addressed the challenge…

Scaling Laws for Optimal Data Mixtures

2025-07-12 · Mustafa Shukor, Louis Bethune, Dan Busbridge, David Grangier 외 arxiv

Large foundation models are typically trained on data from multiple domains, with the data mixture--the proportion of each domain used--playing a critical role in model performance. The standard approach to selecting thi…

Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model

2026-05-14 · Arie Wortsman-Zurich, Hugo Tabanelli, Yatin Dandi, Florent Krzakala 외 arxiv

We propose a simple mechanism by which scaling laws emerge from feature learning in multi-layer networks. We study a high-dimensional hierarchical target that is a globally high-degree function, but that can be represent…

Sharp feature-learning transitions and Bayes-optimal neural scaling laws in extensive-width networks

2026-05-11 · Minh-Toan Nguyen, Jean Barbier arxiv

We study the information-theoretic limits of learning a one-hidden-layer teacher network with hierarchical features from noisy queries, in the context of knowledge transfer to a smaller student model. We work in the high…