paper-with-me

홈 › Papers

Foundational Moral Values for AI Alignment

2023-11-28 · Betty Li Hou, Brian Patrick Green

Solving the AI alignment problem requires having clear, defensible values towards which AI systems can align. Currently, targets for alignment remain underspecified and do not seem to be built from a philosophically robust structure. We begin the discussion of this problem by presenting five core, foundational values, drawn from moral philosophy and built on the requisites for human existence: survival, sustainable intergenerational existence, society, education, and truth. We show that these values not only provide a clearer direction for technical alignment work, but also serve as a framework to highlight threats and opportunities from AI systems to both obtain and sustain these values.

📄 PDF Abstract BibTeX arXiv:2311.17017

Code (0)

등록된 구현이 없습니다.

Tasks

Philosophy

Similar Papers 제목 키워드 기반

Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety

2025-05-31 · Matthew Brophy

As large language models (LLMs) become more powerful and pervasive across society, ensuring these systems are beneficial, safe, and aligned with human values is crucial. Current alignment techniques, like Constitutional …

Brief chatbot interactions produce lasting changes in human moral values

2026-04-23 · Yue Teng, Qianer Zhong, Kim Mai Tich Nguyen Thordsen, Christian Montag 외 arxiv

Moral judgements form the foundation of human social behavior and societal systems. While Artificial Intelligence chatbots increasingly serve as personal advisors, their influence on moral judgments remains largely unexp…

Moral Scenarios

Histoires Morales: A French Dataset for Assessing Moral Alignment

2025-01-28 · Thibaud Leteno, Irina Proskurina, Antoine Gourru, Julien Velcin 외

Aligning language models with human values is crucial, especially as they become more integrated into everyday life. While models are often adapted to user preferences, it is equally important to ensure they align with m…

Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs

2026-01-13 · Chenchen Yuan, Bolei Ma, Zheyu Zhang, Bardh Prenkaj 외 arxiv

While recent research has systematically documented political orientation in large language models (LLMs), existing evaluations rely primarily on direct probing or demographic persona engineering to surface ideological b…

ProgressGym: Alignment with a Millennium of Moral Progress

2024-06-28 · Tianyi Qiu, Yang Zhang, Xuchuan Huang, Jasmine Xinze Li 외

Frontier AI systems, including large language models (LLMs), hold increasing influence over the epistemology of human users. Such influence can reinforce prevailing societal values, potentially contributing to the lock-i…