paper-with-me

Papers

The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge

2026-05-13 · Ryoya Awano, Taiji Suzuki arxiv

Weak-to-strong (W2S) generalization, in which a strong model is fine-tuned on outputs of a weaker, task-specialized model, has been proposed as an approach to aligning superhuman AI systems. Existing theoretical analyses either fix the student's representations or operate in restricted settings. Whether multi-step SGD can succeed in feature learning while preserving diverse pre-trained capabilities remains open. We study W2S in the setting of reward-model learning with two-layer neural networks. The strong model has pre-trained representations organized into low-dimensional subspaces $V_k$, and is fine-tuned under the supervision of a weak model specialized on task $κ$. We prove that the strong model efficiently learns task $κ$, eliciting its pre-trained knowledge while retaining general capabilities. This establishes W2S generalization in the feature-learning regime, in the sense that the strong model acquires the target feature direction through W2S training, rather than having it given a priori. Moreover, W2S preserves pre-trained off-target features, whereas standard supervised fine-tuning causes catastrophic forgetting when off-target feature directions are correlated with the target's. Numerical experiments on synthetic data confirm our theoretical results.

📄 PDF Abstract BibTeX arXiv:2605.12908

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning

2025-10-28 · Junsoo Oh, Jerry Song, Chulhee Yun arxiv

Weak-to-strong generalization refers to the phenomenon where a stronger model trained under supervision from a weaker one can outperform its teacher. While prior studies aim to explain this effect, most theoretical insig…

Weak-to-Strong Generalization Through the Data-Centric Lens

2024-12-05 · Changho Shin, John Cooper, Frederic Sala

The weak-to-strong generalization phenomenon is the driver for important machine learning applications including highly data-efficient learning and, most recently, performing superalignment. While decades of research hav…

Recognizing and Eliciting Weakly Single Crossing Profiles on Trees

2016-11-13 · Palash Dey

We introduce and study the weakly single-crossing domain on trees which is a generalization of the well-studied single-crossing domain in social choice theory. We design a polynomial-time algorithm for recognizing prefer…

Open-Ended Question Answering

On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective

2025-05-23 · Behrad Moniri, Hamed Hassani

Weak-to-strong generalization, where a student model trained on imperfect labels generated by a weaker teacher nonetheless surpasses that teacher, has been widely observed but the mechanisms that enable it have remained …

regression

Mechanisms for belief elicitation without ground truth

2024-09-11 · Niklas Valentin Lehmann

This review article examines the challenge of eliciting truthful information from multiple individuals when such information cannot be verified, a problem known as information elicitation without verification (IEWV). Thi…