paper-with-me

홈 › Papers

Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives

2025-10-30 · Kentaro Ozeki, Risako Ando, Takanobu Morishita, Hirohiko Abe, Koji Mineshima, Mitsuhiro Okada arxiv

Normative reasoning is a type of reasoning that involves normative or deontic modality, such as obligation and permission. While large language models (LLMs) have demonstrated remarkable performance across various reasoning tasks, their ability to handle normative reasoning remains underexplored. In this paper, we systematically evaluate LLMs' reasoning capabilities in the normative domain from both logical and modal perspectives. Specifically, to assess how well LLMs reason with normative modals, we make a comparison between their reasoning with normative modals and their reasoning with epistemic modals, which share a common formal structure. To this end, we introduce a new dataset covering a wide range of formal patterns of reasoning in both normative and epistemic domains, while also incorporating non-formal cognitive factors that influence human reasoning. Our results indicate that, although LLMs generally adhere to valid reasoning patterns, they exhibit notable inconsistencies in specific types of normative reasoning and display cognitive biases similar to those observed in psychological studies of human reasoning. These findings highlight challenges in achieving logical consistency in LLMs' normative reasoning and provide insights for enhancing their reliability. All data and code are released publicly at https://github.com/kmineshima/NeuBAROCO.

📄 PDF Abstract BibTeX arXiv:2510.26606

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bridging between LegalRuleML and TPTP for Automated Normative Reasoning (extended version)

2022-09-12 · Alexander Steen, David Fuenmayor

LegalRuleML is a comprehensive XML-based representation framework for modeling and exchanging normative rules. The TPTP input and output formats, on the other hand, are general-purpose standards for the interaction with …

Translation

Harnessing the power of LLMs for normative reasoning in MASs

2024-03-25 · Bastin Tony Roy Savarimuthu, Surangika Ranathunga, Stephen Cranefield

Software agents, both human and computational, do not exist in isolation and often need to collaborate or coordinate with others to achieve their goals. In human society, social mechanisms such as norms ensure efficient …

Decision Making

EgoNormia: Benchmarking Physical Social Norm Understanding

2025-02-27 · MohammadHossein Rezaei, Yicheng Fu, Phil Cuvin, Caleb Ziems 외

Human activity is moderated by norms. However, machines are often trained without explicit supervision on norm understanding and reasoning, particularly when norms are physically- or socially-grounded. To improve and eva…

Answer GenerationBenchmarkingRAG

World model inspired sarcasm reasoning with large language model agents

2025-12-30 · Keito Inoshita, Shinnosuke Mizuno arxiv

Sarcasm understanding is a challenging problem in natural language processing, as it requires capturing the discrepancy between the surface meaning of an utterance and the speaker's intentions as well as the surrounding …

Sarcasm Detection

The Cognitive Capabilities of Generative AI: A Comparative Analysis with Human Benchmarks

2024-10-09 · Isaac R. Galatzer-Levy, David Munday, Jed McGiffin, Xin Liu 외

There is increasing interest in tracking the capabilities of general intelligence foundation models. This study benchmarks leading large language models and vision language models against human performance on the Wechsle…

Retrieval