paper-with-me

홈 › Papers

MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables

2025-09-15 · Matteo Marcuzzo, Alessandro Zangari, Andrea Albarelli, Jose Camacho-Collados, Mohammad Taher Pilehvar arxiv

As LLMs excel on standard reading comprehension benchmarks, attention is shifting toward evaluating their capacity for complex abstract reasoning and inference. Literature-based benchmarks, with their rich narrative and moral depth, provide a compelling framework for evaluating such deeper comprehension skills. Here, we present MORABLES, a human-verified benchmark built from fables and short stories drawn from historical literature. The main task is structured as multiple-choice questions targeting moral inference, with carefully crafted distractors that challenge models to go beyond shallow, extractive question answering. To further stress-test model robustness, we introduce adversarial variants designed to surface LLM vulnerabilities and shortcuts due to issues such as data contamination. Our findings show that, while larger models outperform smaller ones, they remain susceptible to adversarial manipulation and often rely on superficial patterns rather than true moral reasoning. This brittleness results in significant self-contradiction, with the best models refuting their own answers in roughly 20% of cases depending on the framing of the moral choice. Interestingly, reasoning-enhanced models fail to bridge this gap, suggesting that scale - not reasoning ability - is the primary driver of performance.

📄 PDF Abstract BibTeX arXiv:2509.12371

Code (0)

등록된 구현이 없습니다.

Tasks

Reading ComprehensionQuestion Answering

Similar Papers 제목 키워드 기반

MoralBench: Moral Evaluation of LLMs

2024-06-06 · Jianchao Ji, Yutong Chen, Mingyu Jin, Wujiang Xu 외

In the rapidly evolving field of artificial intelligence, large language models (LLMs) have emerged as powerful tools for a myriad of applications, from natural language processing to decision-making support systems. How…

Ethics

The Convergent Ethics of AI? Analyzing Moral Foundation Priorities in Large Language Models with a Multi-Framework Approach

2025-04-27 · Chad Coleman, W. Russell Neuman, Ali Dasdan, Safinah Ali 외

As large language models (LLMs) are increasingly deployed in consequential decision-making contexts, systematically assessing their ethical reasoning capabilities becomes a critical imperative. This paper introduces the …

BenchmarkingDecision MakingEthicsFairness

Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations

2025-05-26 · Chaoyi Xiang, Chunhua Liu, Simon De Deyne, Lea Frermann

As the impact of large language models increases, understanding the moral values they reflect becomes ever more important. Assessing the nature of moral values as understood by these models via direct prompting is challe…

Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs

2023-10-11 · Abhinav Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal 외

In this position paper, we argue that instead of morally aligning LLMs to specific set of ethical principles, we should infuse generic ethical reasoning capabilities into them so that they can handle value pluralism at a…

EthicsPosition

EMNLP: Educator-role Moral and Normative Large Language Models Profiling

2025-08-21 · Yilin Jiang, Mingzi Zhang, Sheng Jin, Zengyi Yu 외 arxiv

Simulating Professions (SP) enables Large Language Models (LLMs) to emulate professional roles. However, comprehensive psychological and ethical evaluation in these contexts remains lacking. This paper introduces EMNLP, …