paper-with-me

Papers

Alignment Makes Language Models Normative, Not Descriptive

2026-03-17 · Eilam Shapira, Moshe Tennenholtz, Roi Reichart arxiv

Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed human behavior. We compare 120 base-aligned model pairs on more than 10,000 real human decisions in multi-round strategic games - bargaining, persuasion, negotiation, and repeated matrix games. In these settings, base models outperform their aligned counterparts in predicting human choices by nearly 10:1, robustly across model families, prompt formulations, and game configurations. This pattern reverses, however, in settings where human behavior is more likely to follow normative predictions: aligned models dominate on one-shot textbook games across all 12 types tested and on non-strategic lottery choices - and even within the multi-round games themselves, at round one, before interaction history develops. This boundary-condition pattern suggests that alignment induces a normative bias: it improves prediction when human behavior is relatively well captured by normative solutions, but hurts prediction in multi-round strategic settings, where behavior is shaped by descriptive dynamics such as reciprocity, retaliation, and history-dependent adaptation. These results reveal a fundamental trade-off between optimizing models for human use and using them as proxies for human behavior.

📄 PDF Abstract BibTeX arXiv:2603.17218

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Measuring Normative and Descriptive Biases in Language Models Using Census Data

2023-04-12 · Samia Touileb, Lilja Øvrelid, Erik Velldal

We investigate in this paper how distributions of occupations with respect to gender is reflected in pre-trained language models. Such distributions are not always aligned to normative ideals, nor do they necessarily ref…

Descriptive

Beyond Preferences in AI Alignment

2024-08-30 · Tan Zhi-Xuan, Micah Carroll, Matija Franklin, Hal Ashton

The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizing the satisfaction of preferences, and …

Descriptive

Maximin Safety: When Failing to Lose is Preferable to Trying to Win

2015-01-21 · Brad Gulko, Samantha Leung

We present a new decision rule, \emph{maximin safety}, that seeks to maintain a large margin from the worst outcome, in much the same way minimax regret seeks to minimize distance from the best. We argue that maximin saf…

Differences in the Moral Foundations of Large Language Models

2025-11-14 · Peter Kirgis arxiv

Large language models are increasingly being used in critical domains of politics, business, and education, but the nature of their normative ethical judgment remains opaque. Alignment research has, to date, not sufficie…

From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment

2024-06-20 · Yusuke Hirota, Ryo Hachiuma, Chao-Han Huck Yang, Yuta Nakashima

Large language models (LLMs) have enhanced the capacity of vision-language models to caption visual text. This generative approach to image caption enrichment further makes textual captions more descriptive, improving al…

DescriptiveHallucination