paper-with-me

Papers

Superalignment with Dynamic Human Values

2025-03-17 · Florian Mai, David Kaczér, Nicholas Kluge Corrêa, Lucie Flek

Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they do not simultaneously account for 2). We sketch a roadmap for a novel algorithmic framework that trains a superhuman reasoning model to decompose complex tasks into subtasks that are still amenable to human-level guidance. Our approach relies on what we call the part-to-complete generalization hypothesis, which states that the alignment of subtask solutions generalizes to the alignment of complete solutions. We advocate for the need to measure this generalization and propose ways to improve it in the future.

📄 PDF Abstract BibTeX arXiv:2503.13621

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Moral Imperative: The Need for Continual Superalignment of Large Language Models

2024-03-13 · Gokul Puthumanaillam, Manav Vora, Pranay Thangeda, Melkior Ornik

This paper examines the challenges associated with achieving life-long superalignment in AI systems, particularly large language models (LLMs). Superalignment is a theoretical framework that aspires to ensure that superi…

Ethics

The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment

2024-12-21 · HyunJin Kim, Xiaoyuan Yi, Jing Yao, Jianxun Lian 외

The emergence of large language models (LLMs) has sparked the possibility of about Artificial Superintelligence (ASI), a hypothetical AI system surpassing human intelligence. However, existing alignment paradigms struggl…

Survey

Super Co-alignment for Sustainable Symbiotic Society

2025-04-24 · Yi Zeng, Feifei Zhao, Yuwei Wang, Enmeng Lu 외

As Artificial Intelligence (AI) advances toward Artificial General Intelligence (AGI) and eventually Artificial Superintelligence (ASI), it may potentially surpass human control, deviate from human values, and even lead …

Research on Superalignment Should Advance Now with Parallel Optimization of Competence and Conformity

2025-03-08 · HyunJin Kim, Xiaoyuan Yi, Jing Yao, Muhua Huang 외

The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussions on Artificial Superintelligence (ASI),…

The Superalignment of Superhuman Intelligence with Large Language Models

2024-12-15 · Minlie Huang, Yingkang Wang, Shiyao Cui, Pei Ke 외

We have witnessed superhuman intelligence thanks to the fast development of large language models and multimodal language models. As the application of such superhuman models becomes more and more popular, a critical que…