paper-with-me

Papers

Weak-to-Strong Generalization beyond Accuracy: a Pilot Study in Safety, Toxicity, and Legal Reasoning

2024-10-16 · Ruimeng Ye, Yang Xiao, Bo Hui

As large language models (LLMs) continue to advance, ensuring their alignment with human values becomes increasingly critical. Traditional alignment methods heavily rely on human feedback to fine-tune models. With the emergence of superhuman models whose outputs may surpass human understanding, evaluating and aligning these models using human judgments poses significant challenges. To address the challenges, recent works use weak supervisors to elicit knowledge from much stronger models. However, there are important disanalogies between the empirical setup in the existing works and the genuine goal of alignment. We remark that existing works investigate the phenomenon of weak-to-strong generation in analogous setup (i.e., binary classification), rather than practical alignment-relevant tasks (e.g., safety). In this paper, we bridge this gap by extending weak-to-strong generation to the context of practical alignment. We empirically demonstrate the widespread phenomenon of weak-to-strong generation in three complicated alignment tasks: safety, toxicity, and legal reasoning}. Furthermore, we explore efficient strategies for improving alignment performance to enhance the quality of model outcomes. Lastly, we summarize and analyze the challenges and potential solutions in regard to specific alignment tasks, which we hope to catalyze the research progress on the topic of weak-to-strong generalization. Our code is released at https://github.com/yeruimeng/WTS.git.

📄 PDF Abstract BibTeX arXiv:2410.12621

Code (1)

yeruimeng/wts 공식 구현 pytorch

Tasks

Binary ClassificationLegal Reasoning

Similar Papers 제목 키워드 기반

UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization

2026-04-15 · Zhengxi Lu, Fei Tang, Guangyi Liu, Kaitao Song 외 arxiv

MLLM-based GUI agents have demonstrated strong capabilities in complex user interface interaction tasks. However, long-horizon scenarios remain challenging, as these agents are burdened with tasks beyond their intrinsic …

Relating Misfit to Gain in Weak-to-Strong Generalization Beyond the Squared Loss

2025-01-31 · Abhijeet Mulgund, Chirag Pabbaraju

The paradigm of weak-to-strong generalization constitutes the training of a strong AI model on data labeled by a weak AI model, with the goal that the strong model nevertheless outperforms its weak supervisor on the targ…

Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models

2024-02-06 · Jianyuan Guo, Hanting Chen, Chengcheng Wang, Kai Han 외

Recent advancements in large language models have sparked interest in their extraordinary and near-superhuman capabilities, leading researchers to explore methods for evaluating and optimizing these abilities, which is c…

Few-Shot LearningKnowledge DistillationTransfer Learning

Selective Weak-to-Strong Generalization

2025-11-18 · Hao Lang, Fei Huang, Yongbin Li arxiv

Future superhuman models will surpass the ability of humans and humans will only be able to \textit{weakly} supervise superhuman models. To alleviate the issue of lacking high-quality data for model alignment, some works…

Bayesian WeakS-to-Strong from Text Classification to Generation

2024-05-24 · Ziyun Cui, Ziyang Zhang, Wen Wu, Guangzhi Sun 외

Advances in large language models raise the question of how alignment techniques will adapt as models become increasingly complex and humans will only be able to supervise them weakly. Weak-to-Strong mimics such a scenar…

text-classificationText ClassificationText Generation