paper-with-me

홈 › Papers

ProgressGym: Alignment with a Millennium of Moral Progress

2024-06-28 · Tianyi Qiu, Yang Zhang, Xuchuan Huang, Jasmine Xinze Li, Jiaming Ji, Yaodong Yang

Frontier AI systems, including large language models (LLMs), hold increasing influence over the epistemology of human users. Such influence can reinforce prevailing societal values, potentially contributing to the lock-in of misguided moral beliefs and, consequently, the perpetuation of problematic moral practices on a broad scale. We introduce progress alignment as a technical solution to mitigate this imminent risk. Progress alignment algorithms learn to emulate the mechanics of human moral progress, thereby addressing the susceptibility of existing alignment methods to contemporary moral blindspots. To empower research in progress alignment, we introduce ProgressGym, an experimental framework allowing the learning of moral progress mechanics from history, in order to facilitate future progress in real-world moral decisions. Leveraging 9 centuries of historical text and 18 historical LLMs, ProgressGym enables codification of real-world progress alignment challenges into concrete benchmarks. Specifically, we introduce three core challenges: tracking evolving values (PG-Follow), preemptively anticipating moral progress (PG-Predict), and regulating the feedback loop between human and AI value shifts (PG-Coevolve). Alignment methods without a temporal dimension are inapplicable to these tasks. In response, we present lifelong and extrapolative algorithms as baseline methods of progress alignment, and build an open leaderboard soliciting novel algorithms and challenges. The framework and the leaderboard are available at https://github.com/PKU-Alignment/ProgressGym and https://huggingface.co/spaces/PKU-Alignment/ProgressGym-LeaderBoard respectively.

📄 PDF Abstract BibTeX arXiv:2406.20087

Code (1)

pku-alignment/progressgym 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Histoires Morales: A French Dataset for Assessing Moral Alignment

2025-01-28 · Thibaud Leteno, Irina Proskurina, Antoine Gourru, Julien Velcin 외

Aligning language models with human values is crucial, especially as they become more integrated into everyday life. While models are often adapted to user preferences, it is equally important to ensure they align with m…

Bounded Morality: Defining the Space of Moral Computation

2026-04-01 · Max Kanwal, Caryn Tran, Patrick Mineault arxiv

Moral cognition has traditionally been modeled as adherence to fixed ethical theories--deontology, consequentialism, virtue ethics--implemented as static rules or value functions. We propose Bounded Morality, a formal fr…

Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models

2026-03-23 · Aryan Kasat, Smriti Singh, Aman Chadha, Vinija Jain arxiv

Do large language models reason morally, or do they merely sound like they do? We investigate whether LLM responses to moral dilemmas exhibit genuine developmental progression through Kohlberg's stages of moral developme…

New Millennium AI and the Convergence of History

2006-06-19 · Juergen Schmidhuber

Artificial Intelligence (AI) has recently become a real formal science: the new millennium brought the first mathematically sound, asymptotically optimal, universal problem solvers, providing a new, rigorous foundation f…

LLMs grasp morality in concept

2023-11-04 · Mark Pock, Andre Ye, Jared Moore

Work in AI ethics and fairness has made much progress in regulating LLMs to reflect certain values, such as fairness, truth, and diversity. However, it has taken the problem of how LLMs might 'mean' anything at all for g…

DiversityEthicsFairnessPhilosophy+1