paper-with-me

홈 › Papers

Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot Tuning

2026-02-19 · Hien Dang, Pratik Patil, Alessandro Rinaldo arxiv

Self-distillation (SD) is the process of retraining a student on a mixture of ground-truth labels and the teacher's own predictions using the same architecture and training data. Although SD has been empirically shown to often improve generalization, its formal guarantees remain limited. We study SD for ridge regression in unconstrained setting in which the mixing weight $ξ$ may be outside the unit interval. Conditioned on the training data and without any distributional assumptions, we prove that for any squared prediction risk (including out-of-distribution), the optimally mixed student strictly improves upon the ridge teacher for every regularization level $λ> 0$ at which the teacher ridge risk $R(λ)$ is nonstationary (i.e., $R'(λ) \neq 0$). We obtain a closed-form expression for the optimal mixing weight $ξ^\star(λ)$ for any value of $λ$ and show that it obeys the sign rule: $\operatorname{sign}(ξ^\star(λ))=-\operatorname{sign}(R'(λ))$. In particular, $ξ^\star(λ)$ can be negative, which is the case in over-regularized regimes. To quantify the risk improvement due to SD, we derive exact deterministic equivalents for the optimal SD risk in the proportional asymptotics regime (where the sample and feature sizes $n$ and $p$ both diverge but their aspect ratio $p/n$ converges) under general anisotropic covariance and deterministic signals. Our asymptotic analysis extends standard second-order ridge deterministic equivalents to their fourth-order analogs using block linearization, which may be of independent interest. From a practical standpoint, we propose a consistent one-shot tuning method to estimate $ξ^\star$ without grid search, sample splitting, or refitting. Experiments on real-world datasets and pretrained neural network features support our theory and the one-shot tuning method.

📄 PDF Abstract BibTeX arXiv:2602.17565

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-Distillation is Optimal Among Spectral Shrinkage Estimators in Spiked Covariance Models

2026-05-18 · Radu Lecoiu, Debarghya Mukherjee, Pragya Sur arxiv

Self-distillation has emerged as a promising technique for improving model performance in modern machine learning systems. We develop the statistical foundations of self-distillation in spiked covariance models, by intro…

Prediction-Only Distillation in Linear and Logistic Regression

2026-07-16 · Hien Dang, Pratik Patil, Alessandro Rinaldo arxiv

Self-distillation (SD) is typically studied when the student is retrained on the teacher's original training inputs. In many practical deployments, however, the labeled training data are no longer available, and one has …

Self-Distillation for Gaussian Process Regression and Classification

2023-04-05 · Kenneth Borup, Lars Nørvang Andersen

We propose two approaches to extend the notion of knowledge distillation to Gaussian Process Regression (GPR) and Gaussian Process Classification (GPC); data-centric and distribution-centric. The data-centric approach re…

ClassificationGPRKnowledge Distillationregression

Optimal Self-Distillation for Rectified Flow via Linear Probing

2026-07-16 · Saptarshi Roy, Debepsita Mukherjee, Pratik Patil arxiv

Modern generative models are increasingly trained using model-generated signals, creating both opportunities for self-improvement and risks of collapse. We study optimal self-distillation (SD) for rectified flow (RF): gi…

Why Self-Training Helps and Hurts: Denoising vs. Signal Forgetting

2026-02-15 · Mingqi Wu, Archer Y. Yang, Qiang Sun arxiv

Iterative self-training (self-distillation) repeatedly refits a model on pseudo-labels generated by its own predictions. We study this procedure in overparameterized linear regression: an initial estimator is trained on …