paper-with-me

홈 › Papers

Divergence of Empirical Neural Tangent Kernel in Classification Problems

2025-04-15 · Zixiong Yu, Songtao Tian, Guhan Chen

This paper demonstrates that in classification problems, fully connected neural networks (FCNs) and residual neural networks (ResNets) cannot be approximated by kernel logistic regression based on the Neural Tangent Kernel (NTK) under overtraining (i.e., when training time approaches infinity). Specifically, when using the cross-entropy loss, regardless of how large the network width is (as long as it is finite), the empirical NTK diverges from the NTK on the training samples as training time increases. To establish this result, we first demonstrate the strictly positive definiteness of the NTKs for multi-layer FCNs and ResNets. Then, we prove that during training, % with the cross-entropy loss, the neural network parameters diverge if the smallest eigenvalue of the empirical NTK matrix (Gram matrix) with respect to training samples is bounded below by a positive constant. This behavior contrasts sharply with the lazy training regime commonly observed in regression problems. Consequently, using a proof by contradiction, we show that the empirical NTK does not uniformly converge to the NTK across all times on the training samples as the network width increases. We validate our theoretical results through experiments on both synthetic data and the MNIST classification task. This finding implies that NTK theory is not applicable in this context, with significant theoretical implications for understanding neural networks in classification problems.

📄 PDF Abstract BibTeX arXiv:2504.11130

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

NTK 설명 없음
Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…

Similar Papers 제목 키워드 기반

NFT-K: Non-Fungible Tangent Kernels

2021-10-11 · Sina AlEMohammad, Hossein Babaei, CJ Barberan, Naiming Liu 외

Deep neural networks have become essential for numerous applications due to their strong empirical performance such as vision, RL, and classification. Unfortunately, these networks are quite difficult to interpret, and t…

An Empirical Analysis of the Laplace and Neural Tangent Kernels

2022-08-07 · Ronaldas Paulius Lencevicius

The neural tangent kernel is a kernel function defined over the parameter distribution of an infinite width neural network. Despite the impracticality of this limit, the neural tangent kernel has allowed for a more direc…

Gradient Descent can Learn Less Over-parameterized Two-layer Neural Networks on Classification Problems

2019-05-23 · Atsushi Nitanda, Geoffrey Chinot, Taiji Suzuki

Recently, several studies have proven the global convergence and generalization abilities of the gradient descent method for two-layer ReLU networks. Most studies especially focused on the regression problems with the sq…

General ClassificationGeneralization Bounds

Globally Convergent Variational Inference

2025-01-14 · Declan McNamara, Jackson Loper, Jeffrey Regier

In variational inference (VI), an approximation of the posterior distribution is selected from a family of distributions through numerical optimization. With the most common variational objective function, known as the e…

Variational Inference

Local Signal Adaptivity: Provable Feature Learning in Neural Networks Beyond Kernels

2021-12-01 · NeurIPS 2021 12 · Stefani Karp, Ezra Winston, Yuanzhi Li, Aarti Singh

Neural networks have been shown to outperform kernel methods in practice (including neural tangent kernels). Most theoretical explanations of this performance gap focus on learning a complex hypothesis class; in some cas…

image-classificationImage Classification