paper-with-me

홈 › Papers

Multigrade Neural Network Approximation

2026-01-23 · Shijun Zhang, Zuowei Shen, Yuesheng Xu arxiv

We study multigrade deep learning (MGDL) as a principled framework for structured error refinement in deep neural networks. While the approximation power of neural networks is now relatively well understood, training very deep architectures remains challenging due to highly nonconvex and often ill-conditioned optimization landscapes. In contrast, for relatively shallow networks, most notably certain one-hidden-layer ReLU models, training admits convex reformulations with global guarantees under appropriate settings, motivating learning paradigms that improve stability while scaling to depth. MGDL builds on this insight by training deep networks grade by grade: previously learned grades are frozen, and each newly added grade-wise subnetwork is composed on top of the previously learned grades and trained to fit the residual left by the current approximation, yielding a structured and interpretable hierarchical refinement process. We develop an operator-theoretic foundation for MGDL and prove that, for any continuous target function defined on a hypercube, there exists a fixed-width multigrade ReLU scheme whose residuals are pointwise nonincreasing in magnitude and converge uniformly to zero, with strict $L^p$-norm decay at every nontrivial grade for $p\in [1,\infty)$. To the best of our knowledge, this work provides the first rigorous constructive approximation guarantee showing that a grade-wise residual refinement scheme can achieve vanishing error in a fixed-width multigrade ReLU architecture.

📄 PDF Abstract BibTeX arXiv:2601.16884

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Layer-wise Geometric Approximation Rates for Deep Networks

2026-04-22 · Shijun Zhang, Zuowei Shen, Yuesheng Xu arxiv

Depth is widely viewed as a central contributor to the success of deep neural networks, whereas standard neural network approximation theory typically provides guarantees only for the final output and leaves the role of …

Algebra and Geometry of Camera Resectioning

2023-09-07 · Erin Connelly, Timothy Duff, Jessie Loucks-Tavitas

We study algebraic varieties associated with the camera resectioning problem. We characterize these resectioning varieties' multigraded vanishing ideals using Gr\"obner basis techniques. As an application, we derive and …

Free resolutions of function classes via order complexes

2019-09-05 · Justin Chen, Christopher Eur, Greg Yang, Mengyuan Zhang

Function classes are collections of Boolean functions on a finite set, which are fundamental objects of study in theoretical computer science. We study algebraic properties of ideals associated to function classes previo…

Learning Theory

On the Effect of Ranking Axioms on IR Evaluation Metrics

2022-07-04 · Fernando Giner

The study of IR evaluation metrics through axiomatic analysis enables a better understanding of their numerical properties. Some works have modelled the effectiveness of retrieval metrics with axioms that capture desirab…

Retrieval

Comparison of Deep Learning Segmentation and Multigrader-annotated Mandibular Canals of Multicenter CBCT scans

2022-05-27 · Jorma Järnstedt, Jaakko Sahlsten, Joel Jaskari, Kimmo Kaski 외

Deep learning approach has been demonstrated to automatically segment the bilateral mandibular canals from CBCT scans, yet systematic studies of its clinical and technical validation are scarce. To validate the mandibula…

Out-of-Distribution Generalization