paper-with-me

Papers

Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective

2026-04-28 · Hamid Osooli, Kareema Batool, Rick Gentry, Tiasa Singha Roy, Ashwin Gupta, Anirudha Ramesh arxiv

Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak model's blind spots. Understanding such failures requires going beyond aggregate accuracy, since weak-to-strong errors depend not only on whether the strong model disagrees with the weak model, but also on how confidence and uncertainty are distributed across examples. In this work, we analyze weak-to-strong alignment through a bias--variance--covariance lens that connects misfit theory to practical post-training pipelines. We derive a misfit-based upper bound on weak-to-strong population risk and study its empirical components using continuous confidence scores. We evaluate four weak-to-strong pipelines spanning supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), and reinforcement learning from AI feedback (RLAIF) on the PKU-SafeRLHF and HH-RLHF datasets. Using a blind-spot deception metric that isolates cases where the strong model is confidently wrong while the weak model is uncertain, we find that strong-model variance is the quantity most strongly associated with blind-spot deception among the BVC quantities we study. Covariance provides additional but weaker information, indicating that weak--strong dependence matters, but does not by itself explain the observed failures. These results suggest that strong-model variance can serve as an early-warning signal for weak-to-strong deception, while blind-spot evaluation helps distinguish whether failures are inherited from weak supervision or arise in regions of weak-model uncertainty.

📄 PDF Abstract BibTeX arXiv:2604.25077

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

2026-05-28 · Jun Rui Huang, Wang Bill Zhu, Ziyi Liu, Nathanael Fast 외 arxiv

Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but the social dynamics of these interactions can create harms that are not…

What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift

2025-04-28 · Jiamin Chang, Haoyang Li, Hammond Pearce, Ruoxi Sun 외

The growing adoption of artificial intelligence (AI) has amplified concerns about trustworthiness, including integrity, privacy, robustness, and bias. To assess and attribute these threats, we propose ConceptLens, a gene…

AttributeData PoisoningSafety Alignment

Weak-to-Strong Generalization beyond Accuracy: a Pilot Study in Safety, Toxicity, and Legal Reasoning

2024-10-16 · Ruimeng Ye, Yang Xiao, Bo Hui

As large language models (LLMs) continue to advance, ensuring their alignment with human values becomes increasingly critical. Traditional alignment methods heavily rely on human feedback to fine-tune models. With the em…

Binary ClassificationLegal Reasoning

LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences

2025-07-25 · Yusuke Hirota, Boyi Li, Ryo Hachiuma, Yueh-Hua Wu 외 arxiv

Large Vision-Language Models (LVLMs) have transformed image captioning, shifting from concise captions to detailed descriptions. We introduce LOTUS, a leaderboard for evaluating detailed captions, addressing three main g…

Image Captioning

Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models

2024-02-06 · Jianyuan Guo, Hanting Chen, Chengcheng Wang, Kai Han 외

Recent advancements in large language models have sparked interest in their extraordinary and near-superhuman capabilities, leading researchers to explore methods for evaluating and optimizing these abilities, which is c…

Few-Shot LearningKnowledge DistillationTransfer Learning