paper-with-me

Papers

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents

2026-02-13 · Simon Rosen, Siddarth Singh, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Victoria Williams, Benjamin Rosman, Geraud Nangue Tasse, Steven James arxiv

Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection of AI safety, moral philosophy, and cognitive science. We introduce Morality Chains, a novel formalism for representing moral norms as ordered deontic constraints, and MoralityGym, a benchmark of 98 ethical-dilemma problems presented as trolley-dilemma-style Gymnasium environments. By decoupling task-solving from moral evaluation and introducing a novel Morality Metric, MoralityGym allows the integration of insights from psychology and philosophy into the evaluation of norm-sensitive reasoning. Baseline results with Safe RL methods reveal key limitations, underscoring the need for more principled approaches to ethical decision-making. This work provides a foundation for developing AI systems that behave more reliably, transparently, and ethically in complex real-world contexts.

📄 PDF Abstract BibTeX arXiv:2602.13372

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs

2026-02-05 · Rohan Subramanian Thomas, Shikhar Shiromani, Abdullah Chaudhry, Ruizhe Li 외 arxiv

Prompt design significantly impacts the moral competence and safety alignment of large language models (LLMs), yet empirical comparisons remain fragmented across datasets and models.We introduce ProMoral-Bench, a unified…

Prompt Engineering

MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models

2025-05-20 · Xiao Lin, Zhining Liu, Ze Yang, Gaotang Li 외

Warning: This paper contains examples of harmful language and images. Reader discretion is advised. Recently, vision-language models have demonstrated increasing influence in morally sensitive domains such as autonomous …

Autonomous DrivingMultimodal Reasoning

Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants

2025-08-18 · Alessio Galatolo, Luca Alberto Rappuoli, Katie Winkle, Meriem Beloucif arxiv

The recent rise in popularity of large language models (LLMs) has prompted considerable concerns about their moral capabilities. Although considerable effort has been dedicated to aligning LLMs with human moral values, e…

The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making

2024-10-09 · Basile Garcia, Crystal Qian, Stefano Palminteri

As large language models (LLMs) become increasingly integrated into society, their alignment with human morals is crucial. To better understand this alignment, we created a large corpus of human- and LLM-generated respon…

Decision MakingMoral Scenarios

Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation

2026-04-09 · Sophie Wu, Andrew Piper arxiv

Stories are key to transmitting values across cultures, but their interpretation varies across linguistic and cultural contexts. Thus, we introduce multilingual story moral generation as a novel culturally grounded evalu…

Semantic Similarity