paper-with-me

홈 › Papers

Fixing the NTK: From Neural Network Linearizations to Exact Convex Programs

2023-09-26 · NeurIPS 2023 11

Recently, theoretical analyses of deep neural networks have broadly focused on two directions: 1) Providing insight into neural network training by SGD in the limit of infinite hidden-layer width and infinitesimally small learning rate (also known as gradient flow) via the Neural Tangent Kernel (NTK), and 2) Globally optimizing the regularized training objective via cone-constrained convex reformulations of ReLU networks. The latter research direction also yielded an alternative formulation of the ReLU network, called a gated ReLU network, that is globally optimizable via efficient unconstrained convex programs. In this work, we interpret the convex program for this gated ReLU network as a Multiple Kernel Learning (MKL) model with a weighted data masking feature map and establish a connection to the NTK. Specifically, we show that for a particular choice of mask weights that do not depend on the learning targets, this kernel is equivalent to the NTK of the gated ReLU network on the training data. A consequence of this lack of dependence on the targets is that the NTK cannot perform better than the optimal MKL kernel on the training set. By using iterative reweighting, we improve the weights induced by the NTK to obtain the optimal MKL kernel which is equivalent to the solution of the exact convex reformulation of the gated ReLU network. We also provide several numerical simulations corroborating our theory. Additionally, we provide an analysis of the prediction error of the resulting optimal kernel via consistency results for the group lasso.

📄 PDF Abstract BibTeX arXiv:2309.15096

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

NTK 설명 없음
SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

An Incremental Path-Following Splitting Method for Linearly Constrained Nonconvex Nonsmooth Programs

2018-01-30 · Linbo Qiao, Wei Liu, Steven Hoi

The stationary point of Problem 2 is NOT the stationary point of Problem 1. We are sorry and we are working on fixing this error.

Non-iterative rigid 2D/3D point-set registration using semidefinite programming

2015-01-04 · Yuehaw Khoo, Ankur Kapoor

We describe a convex programming framework for pose estimation in 2D/3D point-set registration with unknown point correspondences. We give two mixed-integer nonlinear program (MINP) formulations of the 2D/3D registration…

Pose Estimation

Tea: Program Repair Using Neural Network Based on Program Information Attention Matrix

2021-07-17 · Wenshuo Wang, Chen Wu, Liang Cheng, Yang Zhang

The advance in machine learning (ML)-driven natural language process (NLP) points a promising direction for automatic bug fixing for software programs, as fixing a buggy program can be transformed to a translation task. …

Bug fixingProgram RepairTranslation

If It's Not Buggy, Don't Fix It: On the Dynamics of Iterative Bug-fixing with LLMs

2026-09-09 · Xietao Wang-Lin, Anton Isopoussu, Louis Mahon arxiv

Large language models (LLMs) have become ubiquitous in software development, with LLM-based automated program repair tools increasingly used during code review. In this report, we explore the iterative blind use of LLMs …

Program Repair

Exact Instance Compression for Convex Empirical Risk Minimization via Color Refinement

2026-01-31 · Bryan Zhu, Ziang Chen arxiv

Empirical risk minimization (ERM) can be computationally expensive, with standard solvers scaling poorly even in the convex setting. We propose a novel lossless compression framework for convex ERM based on color refinem…