paper-with-me

Papers

Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models

2024-04-17 · Jan-Philipp Fränken, Kanishk Gandhi, Tori Qiu, Ayesha Khawaja, Noah D. Goodman, Tobias Gerstenberg

As AI systems like language models are increasingly integrated into decision-making processes affecting people's lives, it's critical to ensure that these systems have sound moral reasoning. To test whether they do, we need to develop systematic evaluations. We provide a framework that uses a language model to translate causal graphs that capture key aspects of moral dilemmas into prompt templates. With this framework, we procedurally generated a large and diverse set of moral dilemmas -- the OffTheRails benchmark -- consisting of 50 scenarios and 400 unique test items. We collected moral permissibility and intention judgments from human participants for a subset of our items and compared these judgments to those from two language models (GPT-4 and Claude-2) across eight conditions. We find that moral dilemmas in which the harm is a necessary means (as compared to a side effect) resulted in lower permissibility and higher intention ratings for both participants and language models. The same pattern was observed for evitable versus inevitable harmful outcomes. However, there was no clear effect of whether the harm resulted from an agent's action versus from having omitted to act. We discuss limitations of our prompt generation pipeline and opportunities for improving scenarios to increase the strength of experimental effects.

📄 PDF Abstract BibTeX arXiv:2404.10975

Code (1)

cicl-stanford/moral-evals 공식 구현

Tasks

Decision MakingLanguage ModellingMoral Permissibility

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes

2025-10-18 · Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko 외 arxiv

As AI systems progress, we rely more on them to make decisions with us and for us. To ensure that such decisions are aligned with human values, it is imperative for us to understand not only what decisions they make but …

Moral Scenarios

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents

2026-02-13 · Simon Rosen, Siddarth Singh, Ebenezer Gelo, Helen Sarah Robertson 외 arxiv

Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection of AI safety, moral philosophy, and cognitive science. We introduce Morality …

Quasi-Dilemmas for Artificial Moral Agents

2018-07-06 · Daniel Kasenberg, Vasanth Sarathy, Thomas Arnold, Matthias Scheutz 외

In this paper we describe moral quasi-dilemmas (MQDs): situations similar to moral dilemmas, but in which an agent is unsure whether exploring the plan space or the world may reveal a course of action that satisfies all …

AITA Generating Moral Judgements of the Crowd with Reasoning

2023-10-21 · Osama Bsher, Ameer Sabri

Morality is a fundamental aspect of human behavior and ethics, influencing how we interact with each other and the world around us. When faced with a moral dilemma, a person's ability to make clear moral judgments can be…

EthicsNavigateText Generation

Probing the Moral Development of Large Language Models through Defining Issues Test

2023-09-23 · Kumar Tanmay, Aditi Khandelwal, Utkarsh Agarwal, Monojit Choudhury

In this study, we measure the moral reasoning ability of LLMs using the Defining Issues Test - a psychometric instrument developed for measuring the moral development stage of a person according to the Kohlberg's Cogniti…