paper-with-me

Papers

Second-Order Path Kernel Interpolation Formulas in Machine Learning

2026-06-05 · Jin Guo, Roy Y. He, Jean-Michel Morel arxiv

Understanding how training data shape neural network predictions is a central problem in modern learning theory. In 2020, Pedro Domingos proposed an interpolation formula valid for every model learned by deterministic gradient descent. It expresses the model's prediction as an integral, along the optimization path, of a data-dependent kernel that aligns the model's gradients at the test and training data. Such a first-order characterization remains valid for models trained with batch-based stochastic optimization. In this paper, we develop second-order forms of these interpolation formulas. We show that the leading path-kernel interpolation is supplemented by a curvature-weighted interpolation term. For stochastic gradient descent, an additional sampling-induced component appears, coupling the curvature of the prediction with the covariance of mini-batch gradient noise. We also extend the representation to stochastic gradient descent with momentum, where the interpolation structure is preserved but with the weights modified by a memory-related factor. Moreover, we establish a concentration estimate for the terminal prediction, identifying the fluctuation scale around the expected second-order representation. Together, these results provide a refinement of the path-kernel interpretation of neural network prediction.

📄 PDF Abstract BibTeX arXiv:2606.07495

Code (0)

등록된 구현이 없습니다.

Tasks

Stochastic Optimization

Similar Papers 제목 키워드 기반

On Interpolation Formulas Describing Neural Network Generalization

2026-03-14 · Jin Guo, Roy Y. He, Jean-Michel Morel arxiv

In 2020 Domingos introduced an interpolation formula valid for "every model trained by gradient descent". He concluded that such models behave approximately as kernel machines. In this work, we extend the Domingos formul…

Machine Learning of Space-Fractional Differential Equations

2018-08-02 · Mamikon Gulian, Maziar Raissi, Paris Perdikaris, George Karniadakis

Data-driven discovery of "hidden physics" -- i.e., machine learning of differential equation models underlying observed data -- has recently been approached by embedding the discovery problem into a Gaussian Process regr…

BIG-bench Machine Learningregression

Information-Theoretic Limits for the Matrix Tensor Product

2020-05-22 · Galen Reeves

This paper studies a high-dimensional inference problem involving the matrix tensor product of random matrices. This problem generalizes a number of contemporary data science problems including the spiked matrix models u…

Stochastic Block Model

Scaling Gaussian Process Regression with Full Derivative Observations

2025-05-14 · Daniel Huang

We present a scalable Gaussian Process (GP) method that can fit and predict full derivative observations called DSoftKI. It extends SoftKI, a method that approximates a kernel via softmax interpolation from learned inter…

regression

Graph signal interpolation with Positive Definite Graph Basis Functions

2019-12-10

For the interpolation of graph signals with generalized shifts of a graph basis function (GBF), we introduce the concept of positive definite functions on graphs. This concept merges kernel-based interpolation with spect…