paper-with-me

Papers

Effective Frontiers: A Unification of Neural Scaling Laws

2026-02-01 · Jiaxuan Zou, Zixuan Gong, Ye Su, Huayi Tang, Yong Liu arxiv

Neural scaling laws govern the prediction power-law improvement of test loss with respect to model capacity ($N$), datasize ($D$), and compute ($C$). However, existing theoretical explanations often rely on specific architectures or complex kernel methods, lacking intuitive universality. In this paper, we propose a unified framework that abstracts general learning tasks as the progressive coverage of patterns from a long-tail (Zipfian) distribution. We introduce the Effective Frontier ($k_\star$), a threshold in the pattern rank space that separates learned knowledge from the unlearned tail. We prove that reducible loss is asymptotically determined by the probability mass of the tail a resource-dependent frontier truncation. Based on our framework, we derive the precise scaling laws for $N$, $D$, and $C$, attributing them to capacity, coverage, and optimization bottlenecks, respectively. Furthermore, we unify these mechanisms via a Max-Bottleneck principle, demonstrating that the Kaplan and Chinchilla scaling laws are not contradictory, but equilibrium solutions to the same constrained optimization problem under different active bottlenecks.

📄 PDF Abstract BibTeX arXiv:2602.02593

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scaling New Frontiers: Insights into Large Recommendation Models

2024-12-01 · Wei Guo, Hao Wang, Luankang Zhang, Jin Yao Chin 외

Recommendation systems are essential for filtering data and retrieving relevant information across various applications. Recent advancements have seen these systems incorporate increasingly large embedding tables, scalin…

Recommendation Systems

Generalizing Scaling Laws for Dense and Sparse Large Language Models

2025-08-08 · Md Arafat Hossain, Xingfu Wu, Valerie Taylor, Ali Jannesari arxiv

Despite recent advancements of large language models (LLMs), optimally predicting the model size for LLM pretraining or allocating optimal resources still remains a challenge. Several efforts have addressed the challenge…

Scaling Laws Under the Microscope: Predicting Transformer Performance from Small Scale Experiments

2022-02-13 · Maor Ivgi, Yair Carmon, Jonathan Berant

Neural scaling laws define a predictable relationship between a model's parameter count and its performance after training in the form of a power law. However, most research to date has not explicitly investigated whethe…

Model Selection

gzip Predicts Data-dependent Scaling Laws

2024-05-26 · Rohan Pandey

Past work has established scaling laws that predict the performance of a neural language model (LM) as a function of its parameter count and the number of tokens it's trained on, enabling optimal allocation of a fixed co…

Language ModelingLanguage Modelling

Coverage Analysis and Scaling Laws of Ultra-Dense Networks

2020-06-30 · Imene Trigui, Sofiene Affes, Marco Di Renzo, Dushantha Nalin K. Jayakody

In this paper, we develop an innovative approach to quantitatively characterize the performance of ultra-dense wireless networks in a plethora of propagation environments. The proposed framework has the potential of sign…