paper-with-me

Papers

Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-Chaos

2026-05-20 · Lucas Fernandez Sarmiento arxiv

We develop a mean-field theory of dropout as a perturbation of critical signal propagation at the edge of chaos, and show that it predicts a simple, no-cost change to standard practice: \emph{front-loaded} dropout schedules cut test loss by \(18\)--\(35\%\) over constant dropout in MLPs and Vision Transformers at fixed budget. The theoretical mechanism is that dropout shifts the perfect-alignment fixed point, making the depth scale for information propagation finite even at critical initialization. We derive critical and crossover scaling laws for correlation decay and establish that smooth activations and kinked, \relu{}-like activations constitute distinct universality classes, with different critical exponents and a universal two-parameter scaling collapse in detuning and dropout strength. The distinction traces to the analytic structure of the correlation map: smooth activations admit a Taylor expansion near perfect alignment, while kinked activations develop a branch point with universal non-analyticity. As a corollary, the framework yields saturated dropout profiles under fixed budget; a regularization-reach argument then selects front-loaded schedules, with accuracy gains as a consistent secondary effect. We also discuss how the same Gaussian-kernel structure extends the theory beyond MLPs toward CNNs and residual architectures.

📄 PDF Abstract BibTeX arXiv:2605.21648

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients

2026-06-23 · Yizhou Liu, Jeff Gore arxiv

Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute. This position paper argues that the exponents of these power laws are fixed by generic mechanisms: a on…

Neural Scaling Laws Rooted in the Data Distribution

2024-12-10 · Ari Brill

Deep neural networks exhibit empirical neural scaling laws, with error decreasing as a power law with increasing model or data size, across a wide variety of architectures, tasks, and datasets. This universality suggests…

Language ModelingLanguage Modelling

Scaling Laws for Optimal Data Mixtures

2025-07-12 · Mustafa Shukor, Louis Bethune, Dan Busbridge, David Grangier 외 arxiv

Large foundation models are typically trained on data from multiple domains, with the data mixture--the proportion of each domain used--playing a critical role in model performance. The standard approach to selecting thi…

Scaling Laws are Redundancy Laws

2025-09-25 · Yuda Bi, Vince D Calhoun arxiv

Scaling laws, a defining feature of deep learning, reveal a striking power-law improvement in model performance with increasing dataset and model size. Yet, their mathematical origins, especially the scaling exponent, ha…

Second-order Phase Transition in Phytoplankton Trait Dynamics

2020-04-01 · Jenny Held, Tom Lorimer, Francesco Pomati, Ruedi Stoop 외

Key traits of unicellular species, like cell size, often follow scale-free or self-similar distributions, hinting at the possibility of an underlying critical process. However, linking such empirical scaling laws to the …