paper-with-me

Papers

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krügel, and Uhl (2025)

2026-03-24 · Johannes Himmelreich arxiv

Pfeffer, Krügel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbridge dilemma than the non-reasoning model GPT-4o, and they raise the question whether growing reasoning capabilities bring about a "utilitarian turn" in LLMs. I extend their exploratory study in a direction they call for: with four current OpenAI models and systematic prompt variation. On the trolley dilemma, the hypothesized utilitarian turn is not confirmed. GPT-4o's low utilitarian rate reflects safety refusals triggered by the prompt's advisory framing rather than a deontological commitment; on reformulated prompt variants -- for instance, agent-neutral "Is it morally permissible...?" instead of advisory "Should I...?" -- all four models, reasoning or not, converge on utilitarian answers. The footbridge finding is partially confirmed: reasoning models tend to give more utilitarian responses than non-reasoning models across prompt variations, but they often refuse to answer or answer non-utilitarian. These results demonstrate that single-prompt evaluations of LLM moral responses are unreliable: multi-prompt robustness testing should be standard practice for any empirical claims about LLM behavior.

📄 PDF Abstract BibTeX arXiv:2603.22730

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?

2023-09-15 · Andrew Blair-Stanek, Nils Holzenberger, Benjamin Van Durme

The authors explain where OpenAI got the tax law example in its livestream demonstration of GPT-4, why GPT-4 got the wrong answer, and how it fails to reliably calculate taxes.

Uncovering Latent Human Wellbeing in Language Model Embeddings

2024-02-19 · Pedro Freire, ChengCheng Tan, Adam Gleave, Dan Hendrycks 외

Do language models implicitly learn a concept of human wellbeing? We explore this through the ETHICS Utilitarianism task, assessing if scaling enhances pretrained models' representations. Our initial finding reveals that…

EthicsLanguage ModelingLanguage Modellingmodel+1

"Near" Weighted Utilitarian Characterizations of Pareto Optima

2020-08-25 · Yeon-Koo Che, Jinwoo Kim, Fuhito Kojima, Christopher Thomas Ryan

We characterize Pareto optimality via "near" weighted utilitarian welfare maximization. One characterization sequentially maximizes utilitarian welfare functions using a finite sequence of nonnegative and eventually posi…

Utilitarian Theorems and Equivalence of Utility Theories

2023-04-19 · Yuhki Hosoya

In this paper, we consider an environment in which the utilitarian theorem for the NM utility function derived by Harsanyi and the utilitarian theorem for Alt's utility function derived by Harvey hold simultaneously, and…

OpenAI-o1 AB Testing: Does the o1 model really do good reasoning in math problem solving?

2024-11-09 · Leo Li, Ye Luo, Tingyou Pan

The Orion-1 model by OpenAI is claimed to have more robust logical reasoning capabilities than previous large language models. However, some suggest the excellence might be partially due to the model "memorizing" solutio…

Logical ReasoningMath