paper-with-me

홈 › Papers

Weak-to-Strong Generalization is Nearly Inevitable (in Linear Models)

2026-05-07 · Scott Geng, Dutch Hansen, Jerry Li arxiv

Weak-to-strong generalization is a phenomenon in post-training whereby a strong student model, when finetuned solely with feedback from a weaker teacher, can not only surpass the teacher, but can improve upon its own capabilities. Recent work of Burns et al. (2023) demonstrated that this can occur in the setting of frontier language models, and subsequently there has been a flurry of both empirical work trying to exploit this phenomenon, as well as theoretical work attempting to understand it. In this work, we demonstrate that weak-to-strong generalization occurs in standard linear logistic regression, under mild distributional assumptions on the data. In fact, we show that this happens for most student-teacher pairs, suggesting that weak-to-strong generalization is in fact \emph{almost inevitable}, even in this basic setting. Notably, our setting does not require the student to be more expressive or have more model capacity in any way compared to the teacher, which runs contrary to the prevailing theoretical belief that a mismatch in model capacity is a central mechanism to weak-to-strong generalization.

📄 PDF Abstract BibTeX arXiv:2605.05742

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning

2025-10-28 · Junsoo Oh, Jerry Song, Chulhee Yun arxiv

Weak-to-strong generalization refers to the phenomenon where a stronger model trained under supervision from a weaker one can outperform its teacher. While prior studies aim to explain this effect, most theoretical insig…

Linear Convergence of the Randomized Feasible Descent Method Under the Weak Strong Convexity Assumption

2015-06-08 · Chenxin Ma, Rachael Tappenden, Martin Takáč

In this paper we generalize the framework of the feasible descent method (FDM) to a randomized (R-FDM) and a coordinate-wise random feasible descent method (RC-FDM) framework. We show that the famous SDCA algorithm for o…

Bandit Multiclass Linear Classification for the Group Linear Separable Case

2019-12-21 · Jittat Fakcharoenphol, Chayutpong Prompak

We consider the online multiclass linear classification under the bandit feedback setting. Beygelzimer, P\'{a}l, Sz\"{o}r\'{e}nyi, Thiruvenkatachari, Wei, and Zhang [ICML'19] considered two notions of linear separability…

ClassificationGeneral Classification

Performance of Empirical Risk Minimization For Principal Component Regression

2024-09-05 · Christian Brownlees, Guðmundur Stefán Guðmundsson, Yaping Wang

This paper establishes bounds on the predictive performance of empirical risk minimization for principal component regression. Our analysis is nonparametric, in the sense that the relation between the prediction target a…

Predictionregression

Learning-based solutions to nonlinear hyperbolic PDEs: Empirical insights on generalization errors

2023-02-16 · Bilal Thonnam Thodi, Sai Venkata Ramana Ambadipudi, Saif Eddin Jabari

We study learning weak solutions to nonlinear hyperbolic partial differential equations (H-PDE), which have been difficult to learn due to discontinuities in their solutions. We use a physics-informed variant of the Four…