paper-with-me

홈 › Papers

A Qualitative Test-Risk Mechanism for Scaling Behavior in Normalized Residual Networks

2026-05-08 · Daning Cheng, Zeyu Liu, Jun Sun, Fen Xia, Boyang Zhang, Dongping Liu, Yunquan Zhang arxiv

The scaling behavior, in which test performance often improves as model size and data increase, is a central empirical phenomenon in modern deep learning, yet its theoretical basis remains incomplete. In this paper, we study depth expansion in normalized residual networks: starting from a trained model in an old hypothesis class, we insert a new residual block at an intermediate layer and ask when such an expansion can yield a provable improvement in test risk. We develop a unified framework that decomposes this question into representational gain, optimization gain, and generalization transfer. First, under a first-order descent condition near zero initialization, we prove that the expanded hypothesis class contains an auxiliary jumpboard model with strictly smaller population risk than the original model. Second, under norm control tailored to post-normalized residual architectures, we establish a norm-based Rademacher complexity bound for the expanded model class. These ingredients lead to two complementary test-risk guarantees: one route passes through population risk and is tighter when a positive population margin is available, while the other works directly at the train/test level, avoids Hoeffding transfer, and is more robust in degenerate regimes. Together, these results provide a theorem-driven mechanism under which residual depth expansion can improve test performance in normalized residual networks. More broadly, they suggest that scaling is inherently joint: depth creates new improving directions, width enhances the finite-sample observability of weak signals, and data determines whether the statistical cost of expansion can be controlled.

📄 PDF Abstract BibTeX arXiv:2605.08297

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the statistics of scaling exponents and the Multiscaling Value at Risk

2020-02-11 · Giuseppe Brandi, T. Di Matteo

Scaling and multiscaling financial time series have been widely studied in the literature. The research on this topic is vast and still flourishing. One way to analyze the scaling properties of time series is through the…

Time SeriesTime Series Analysis

Muse Spark Safety & Preparedness Report

2026-05-14 · Cristina Menghini, Peter Ney, Hamza Kwisaba, Zifan 외 arxiv

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framework, along with the evidence that informe…

When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs

2025-10-25 · Keyu Wang, Tian Lyu, Guinan Su, Jonas Geiping 외 arxiv

Layer pruning has emerged as a widely adopted technique for improving the efficiency of large language models (LLMs). Although existing methods demonstrate strong performance retention on general knowledge tasks, their e…

General Knowledge

Paradox resolved: The allometric scaling of cancer risk across species

2020-11-22 · Christopher P. Kempes, Geoffrey B. West, John W. Pepper

Understanding the cross-species behavior of cancer is important for uncovering fundamental mechanisms of carcinogenesis, and for translating results of model systems between species. One of the most famous interspecific …

An investigation of higher order moments of empirical financial data and the implications to risk

2021-03-24 · Luke De Clerk, Sergey Savel'ev

Here, we analyse the behaviour of the higher order standardised moments of financial time series when we truncate a large data set into smaller and smaller subsets, referred to below as time windows. We look at the effec…

Time SeriesTime Series Analysis