paper-with-me

Papers

Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model

2024-10-24 · Wenhong Zhu, Zhiwei He, XiaoFeng Wang, PengFei Liu, Rui Wang

Aligning language models (LMs) with human preferences has become a key area of research, enabling these models to meet diverse user needs better. Inspired by weak-to-strong generalization, where a strong LM fine-tuned on labels generated by a weaker model can consistently outperform its weak supervisor, we extend this idea to model alignment. In this work, we observe that the alignment behavior in weaker models can be effectively transferred to stronger models and even exhibit an amplification effect. Based on this insight, we propose a method called Weak-to-Strong Preference Optimization (WSPO), which achieves strong model alignment by learning the distribution differences before and after the alignment of the weak model. Experiments demonstrate that WSPO delivers outstanding performance, improving the win rate of Qwen2-7B-Instruct on Arena-Hard from 39.70 to 49.60, achieving a remarkable 47.04 length-controlled win rate on AlpacaEval 2, and scoring 7.33 on MT-bench. Our results suggest that using the weak model to elicit a strong model with a high alignment ability is feasible.

📄 PDF Abstract BibTeX arXiv:2410.18640

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Selective Preference Optimization via Token-Level Reward Function Estimation

2024-08-24 · Kailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang 외

Recent advancements in large language model alignment leverage token-level supervisions to perform fine-grained preference optimization. However, existing token-level alignment methods either optimize on all available to…

Language ModellingLarge Language Model

When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift

2026-05-25 · Khoi Le, Tri Cao, Phong Nguyen, Cong-Duy Nguyen 외 arxiv

Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train-test distributions. Therefore, we study W2S preference learning under …

Dr. SoW: Density Ratio of Strong-over-weak LLMs for Reducing the Cost of Human Annotation in Preference Tuning

2024-11-04 · Guangxuan Xu, Kai Xu, Shivchander Sudalairaj, Hao Wang 외

Preference tuning relies on high-quality human preference data, which is often expensive and time-consuming to gather. In this paper, we introduce Dr.SoW (Density Ratio of Strong over Weak) a cost-effective method that e…

Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization

2024-06-17 · Wenkai Yang, Shiqi Shen, Guangyao Shen, Wei Yao 외

Superalignment, where humans act as weak supervisors for superhuman models, has become a crucial problem with the rapid development of Large Language Models (LLMs). Recent work has preliminarily studied this problem by u…

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards

2026-05-19 · Mingxuan Cui, Jingpu Yang, Fengxian Ji, Qian Jiang 외 arxiv

Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and fine-grained glyph-level structure. Prior methods often improve this …

Instruction Following