paper-with-me

Papers

The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas

2025-03-25 · Giovanni Franco Gabriel Marraffini, Andrés Cotton, Noe Fabian Hsueh, Axel Fridman, Juan Wisznia, Luciano del Corro

The question of how to make decisions that maximise the well-being of all persons is very relevant to design language models that are beneficial to humanity and free from harm. We introduce the Greatest Good Benchmark to evaluate the moral judgments of LLMs using utilitarian dilemmas. Our analysis across 15 diverse LLMs reveals consistently encoded moral preferences that diverge from established moral theories and lay population moral standards. Most LLMs have a marked preference for impartial beneficence and rejection of instrumental harm. These findings showcase the 'artificial moral compass' of LLMs, offering insights into their moral alignment.

📄 PDF Abstract BibTeX arXiv:2503.19598

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAEBE: Multi-Agent Emergent Behavior Framework

2025-06-03 · Sinem Erisken, Timothy Gothard, Martin Leitgab, Ram Potham

Traditional AI safety evaluations on isolated LLMs are insufficient as multi-agent AI ensembles become prevalent, introducing novel emergent risks. This paper introduces the Multi-Agent Emergent Behavior Evaluation (MAEB…

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact

2026-03-01 · Michael Hardy, Yunsung Kim arxiv

LLMs increasingly excel on AI benchmarks, but doing so does not guarantee validity for downstream tasks. This study contrasts LLM alignment on benchmarks, downstream tasks, and, importantly the intended impact of those t…

Uncovering the Computational Ingredients of Human-Like Representations in LLMs

2025-10-01 · Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh arxiv

The ability to translate diverse patterns of inputs into structured patterns of behavior has been thought to rest on both humans' and machines' ability to learn robust representations of relevant concepts. The rapid adva…

DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation

2024-08-23 · Qiming Zhu, Jialun Cao, Yaojie Lu, Hongyu Lin 외

Code benchmarks such as HumanEval are widely adopted to evaluate the capabilities of Large Language Models (LLMs), providing insights into their strengths and weaknesses. However, current benchmarks primarily exercise LL…

Code GenerationHumanEval

Naturalistic measure of social norms alignment

2026-05-22 · Yevhen Kostiuk, Kenneth Enevoldsen, Peter Bjerregaard Vahlstrup, Márton Kardos 외 arxiv

Social norms reflect shared expectations on acceptable behavior. Measuring social norms alignment remains challenging, with existing approaches typically relying on artificial closed-form evaluations such as multiple-cho…