paper-with-me

홈 › Papers

Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability

2026-03-16 · Fan Huang, Haewoon Kwak, Jisun An arxiv

Large language models (LLMs) increasingly participate in morally sensitive decision-making, yet how they organize ethical frameworks across reasoning steps remains underexplored. We introduce \textit{moral reasoning trajectories}, sequences of ethical framework invocations across intermediate reasoning steps, and analyze their dynamics across six models and three benchmarks. We find that moral reasoning involves systematic multi-framework deliberation: 55.4--57.7\% of consecutive steps involve framework switches, and only 16.4--17.8\% of trajectories remain framework-consistent. Unstable trajectories remain 1.29$\times$ more susceptible to persuasive attacks ($p=0.015$). At the representation level, linear probes localize framework-specific encoding to model-specific layers (layer 63/81 for Llama-3.3-70B; layer 17/81 for Qwen2.5-72B), achieving 13.8--22.6\% lower KL divergence than the training-set prior baseline. Lightweight activation steering modulates framework integration patterns (6.7--8.9\% drift reduction) and amplifies the stability--accuracy relationship. We further propose a Moral Representation Consistency (MRC) metric that correlates strongly ($r=0.715$, $p<0.0001$) with LLM coherence ratings, whose underlying framework attributions are validated by human annotators (mean cosine similarity $= 0.859$).

📄 PDF Abstract BibTeX arXiv:2603.16017

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral

2025-02-19 · Shivani Kumar, David Jurgens

Moral reasoning is a complex cognitive process shaped by individual experiences and cultural contexts and presents unique challenges for computational analysis. While natural language processing (NLP) offers promising to…

Probing the Moral Development of Large Language Models through Defining Issues Test

2023-09-23 · Kumar Tanmay, Aditi Khandelwal, Utkarsh Agarwal, Monojit Choudhury

In this study, we measure the moral reasoning ability of LLMs using the Defining Issues Test - a psychometric instrument developed for measuring the moral development stage of a person according to the Kohlberg's Cogniti…

Let's Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning

2023-06-25 · Xiao Ma, Swaroop Mishra, Ahmad Beirami, Alex Beutel 외

Language models still struggle on moral reasoning, despite their impressive performance in many other tasks. In particular, the Moral Scenarios task in MMLU (Multi-task Language Understanding) is among the worst performi…

counterfactualMathMMLUMoral Scenarios+1

Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations

2025-05-26 · Chaoyi Xiang, Chunhua Liu, Simon De Deyne, Lea Frermann

As the impact of large language models increases, understanding the moral values they reflect becomes ever more important. Assessing the nature of moral values as understood by these models via direct prompting is challe…

Common Sense vs. Morality: The Curious Case of Narrative Focus Bias in LLMs

2026-03-10 · Saugata Purkayastha, Pranav Kushare, Pragya Paramita Pal, Sukannya Purkayastha arxiv

Large Language Models (LLMs) are increasingly deployed across diverse real-world applications and user communities. As such, it is crucial that these models remain both morally grounded and knowledge-aware. In this work,…