paper-with-me

홈 › Papers

Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime

2026-05-24 · Ruchirinkil Marreddy, Chaoyue Liu arxiv

The Neural Tangent Kernel (NTK) framework explains optimization in over-parameterized neural networks via approximately linearized dynamics, yielding exponential convergence guarantees. However, existing results are often overly pessimistic and do not match the fast training in practice, as they depend on the smallest NTK eigenvalue, which is typically extremely small in practice. In this work, we develop sharper convergence guarantees by characterizing the interaction between data labels and the NTK eigen-spectrum. We identify two key phenomena, Label-NTK alignment and Residual-NTK alignment, showing that projections of labels and residuals onto NTK eigenvectors scale with the corresponding eigenvalues. We provide empirical evidence and theoretical justification under mild data assumptions. Exploiting these alignment properties, we derive a refined convergence bound that depends on the full spectrum and closely matches practical training dynamics, significantly improving over classical worst-case results. We further obtain improved generalization bounds. Experiments on MLPs and CNNs across multiple datasets validate our theory.

📄 PDF Abstract BibTeX arXiv:2605.25275

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Non-Asymptotic Optimization and Generalization Bounds for Stochastic Gauss-Newton in Overparameterized Models

2025-11-06 · Semih Cayci arxiv

An important question in deep learning is how higher-order optimization methods affect generalization. In this work, we analyze a stochastic Gauss-Newton (SGN) method with Levenberg-Marquardt damping and mini-batch sampl…

Faster Rates For Federated Variational Inequalities

2026-02-09 · Guanghui Wang, Satyen Kale arxiv

In this paper, we study federated optimization for solving stochastic variational inequalities (VIs), a problem that has attracted growing attention in recent years. Despite substantial progress, a significant gap remain…

Convergence of AdaGrad for Non-convex Objectives: Simple Proofs and Relaxed Assumptions

2023-05-29 · Bohan Wang, Huishuai Zhang, Zhi-Ming Ma, Wei Chen

We provide a simple convergence proof for AdaGrad optimizing non-convex objectives under only affine noise variance and bounded smoothness assumptions. The proof is essentially based on a novel auxiliary function $\xi$ t…

Optimal Inference in Crowdsourced Classification via Belief Propagation

2016-02-11 · Jungseul Ok, Sewoong Oh, Jinwoo Shin, Yung Yi

Crowdsourcing systems are popular for solving large-scale labelling tasks with low-paid workers. We study the problem of recovering the true labels from the possibly erroneous crowdsourced labels under the popular Dawid-…

ClassificationGeneral Classification

Tighter 'uniform bounds for Black-Scholes implied volatility' and the applications to root-finding

2023-02-17 · Jaehyuk Choi, Jeonggyu Huh, Nan Su

Using the option delta systematically, we derive tighter lower and upper bounds of the Black-Scholes implied volatility than those in Tehranchi [SIAM J. Financ. Math. 7 (2016), 893-916]. As an application, we propose a N…

Math