paper-with-me

Papers

How Language Models Organize and Structure Moral Knowledge

2026-08-27 · Orion Reblitz-Richardson arxiv

How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.

📄 PDF Abstract BibTeX arXiv:2608.27402

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tracing Moral Foundations in Large Language Models

2026-01-09 · Chenxiao Yu, Bowen Yi, Farzan Karimi-Malekabadi, Suhaib Abdurahman 외 arxiv

Large language models often produce human-like moral judgments, but it is unclear whether this reflects an internal conceptual structure or superficial ``moral mimicry.'' Using Moral Foundations Theory (MFT) as an analyt…

MOKA: Moral Knowledge Augmentation for Moral Event Extraction

2023-11-16 · Xinliang Frederick Zhang, Winston Wu, Nick Beauchamp, Lu Wang

News media often strive to minimize explicit moral language in news articles, yet most articles are dense with moral values as expressed through the reported events themselves. However, values that are reflected in the i…

ArticlesEvent ExtractionMoral Scenarios

Knowledge of cultural moral norms in large language models

2023-06-02 · Aida Ramezani, Yang Xu

Moral norms vary across cultures. A recent line of work suggests that English large language models contain human-like moral biases, but these studies typically do not examine moral variation in a diverse cultural settin…

DiversitySurvey

The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems

2022-04-06 · ACL 2022 5 · Caleb Ziems, Jane A. Yu, Yi-Chia Wang, Alon Halevy 외

Conversational agents have come increasingly closer to human competence in open-domain dialogue settings; however, such models can reflect insensitive, hurtful, or entirely incoherent viewpoints that erode a user's trust…

AttributeBenchmarking

STREAM: Social data and knowledge collective intelligence platform for TRaining Ethical AI Models

2023-10-09 · Yuwei Wang, Enmeng Lu, Zizhe Ruan, Yao Liang 외

This paper presents Social data and knowledge collective intelligence platform for TRaining Ethical AI Models (STREAM) to address the challenge of aligning AI models with human moral values, and to provide ethics dataset…

Ethics