paper-with-me

홈 › Papers

Truncated Kernel Stochastic Gradient Descent with General Losses and Spherical Radial Basis Functions

2025-10-05 · Jinhui Bai, Andreas Christmann, Lei Shi arxiv

In this paper, we propose a novel kernel stochastic gradient descent (SGD) algorithm for large-scale supervised learning with general losses. Compared to traditional kernel SGD, our algorithm improves efficiency and scalability through an innovative regularization strategy. By leveraging the infinite series expansion of spherical radial basis functions, this strategy projects the stochastic gradient onto a finite-dimensional hypothesis space, which is adaptively scaled according to the bias-variance trade-off, thereby enhancing generalization performance. Based on a new estimation of the spectral structure of the kernel-induced covariance operator, we develop an analytical framework that unifies optimization and generalization analyses. We prove that both the last iterate and the suffix average converge at minimax-optimal rates, and we further establish optimal strong convergence in the reproducing kernel Hilbert space. Our framework accommodates a broad class of classical loss functions, including least-squares, Huber, and logistic losses. Moreover, the proposed algorithm significantly reduces computational complexity and achieves optimal storage complexity by incorporating coordinate-wise updates from linear SGD, thereby avoiding the costly pairwise operations typical of kernel SGD and enabling efficient processing of streaming data. Finally, extensive numerical experiments demonstrate the efficiency of our approach.

📄 PDF Abstract BibTeX arXiv:2510.04237

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Truncated Kernel Stochastic Gradient Descent on Spheres

2024-10-02 · JinHui Bai, Lei Shi

Inspired by the structure of spherical harmonics, we propose the truncated kernel stochastic gradient descent (T-kernel SGD) algorithm with a least-square loss function for spherical data fitting. T-kernel SGD employs a …

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

2026-06-04 · Junyu Zhou, Puyu Wang, Yunwen Lei, Marius Kloft 외 arxiv

Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning. We establish qu…

The Directional Bias Helps Stochastic Gradient Descent to Generalize in Kernel Regression Models

2022-04-29 · Yiling Luo, Xiaoming Huo, Yajun Mei

We study the Stochastic Gradient Descent (SGD) algorithm in nonparametric statistics: kernel regression in particular. The directional bias property of SGD, which is known in the linear regression setting, is generalized…

regression

Noisy Truncated SGD: Optimization and Generalization

2021-02-26 · Yingxue Zhou, Xinyan Li, Arindam Banerjee

Recent empirical work on stochastic gradient descent (SGD) applied to over-parameterized deep learning has shown that most gradient components over epochs are quite small. Inspired by such observations, we rigorously stu…

Ridgeless Regression with Random Features

2022-05-01 · Jian Li, Yong liu, Yingying Zhang

Recent theoretical studies illustrated that kernel ridgeless regression can guarantee good generalization ability without an explicit regularization. In this paper, we investigate the statistical properties of ridgeless …

regression