paper-with-me

Papers

MoralReason: Generalizable Moral Decision Alignment For LLM Agents Using Reasoning-Level Reinforcement Learning

2025-11-15 · Zhiyu An, Wan Du arxiv

Large language models are increasingly influencing human moral decisions, yet current approaches focus primarily on evaluating rather than actively steering their moral decisions. We formulate this as an out-of-distribution moral alignment problem, where LLM agents must learn to apply consistent moral reasoning frameworks to scenarios beyond their training distribution. We introduce Moral-Reason-QA, a novel dataset extending 680 human-annotated, high-ambiguity moral scenarios with framework-specific reasoning traces across utilitarian, deontological, and virtue ethics, enabling systematic evaluation of moral generalization in realistic decision contexts. Our learning approach employs Group Relative Policy Optimization with composite rewards that simultaneously optimize decision alignment and framework-specific reasoning processes to facilitate learning of the underlying moral frameworks. Experimental results demonstrate successful generalization to unseen moral scenarios, with softmax-normalized alignment scores improving by +0.757 for utilitarian and +0.450 for deontological frameworks when tested on out-of-distribution evaluation sets. The experiments also reveal training challenges and promising directions that inform future research. These findings establish that LLM agents can be systematically trained to internalize and apply specific moral frameworks to novel situations, providing a critical foundation for AI safety as language models become more integrated into human decision-making processes.

📄 PDF Abstract BibTeX arXiv:2511.12271

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMoral Scenarios

Similar Papers 제목 키워드 기반

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents

2026-02-13 · Simon Rosen, Siddarth Singh, Ebenezer Gelo, Helen Sarah Robertson 외 arxiv

Evaluating moral alignment in agents navigating conflicting, hierarchically structured human norms is a critical challenge at the intersection of AI safety, moral philosophy, and cognitive science. We introduce Morality …

Moral Alignment for LLM Agents

2024-10-02 · Elizaveta Tennant, Stephen Hailes, Mirco Musolesi

Decision-making agents based on pre-trained Large Language Models (LLMs) are increasingly being deployed across various domains of human activity. While their applications are currently rather specialized, several resear…

Ethics

Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment

2025-11-17 · Jea Kwon, Luiz Felipe Vecchietti, Sungwon Park, Meeyoung Cha arxiv

Humans display significant uncertainty when confronted with moral dilemmas, yet the extent of such uncertainty in machines and AI agents remains underexplored. Recent studies have confirmed the overly confident tendencie…

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

2025-05-25 · Steffen Backmann, David Guzman Piedrahita, Emanuel Tewolde, Rada Mihalcea 외

Recent advances in large language models (LLMs) have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a key AI safety concern. While prior work h…

EthicsNavigate

Visual Distraction Undermines Moral Reasoning in Vision-Language Models

2026-03-17 · Xinyi Yang, Chenheng Xu, Weijun Hong, Ce Mo 외 arxiv

Moral reasoning is fundamental to safe Artificial Intelligence (AI), yet ensuring its consistency across modalities becomes critical as AI systems evolve from text-based assistants to embodied agents. Current safety tech…