paper-with-me

홈 › Papers

Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs

2026-06-10 · Elizaveta Tennant, Benjamin Henke, Anita Keshmirian, Murray Shanahan, Verena Rieser, Kristian Lum, Sydney Levine, Julia Haas arxiv

As LLMs increasingly serve in advisory and deliberative roles, users rely on them for non-verifiable reasoning in domains lacking objective ground truths. However, traditional evaluations of LLM reasoning focus almost exclusively on fact-based domains, such as mathematics and science, leaving uncertainty over whether and to what degree models can handle ambiguous, subjective, or value-laden problems over time. To address this concern, we propose moral reasoning as a paradigmatic subdomain of non-verifiable reasoning. We define moral robustness as a model's capacity to exhibit sound moral reasoning across time and contexts, and we introduce a scalable, adversarial, multi-turn evaluation framework to empirically measure this capability. We simulate 48,000 user-agent moral deliberations across four frontier LLMs, varying premise relevance, premise order, conversation duration, and the user's stated moral view. We find that models successfully ignore morally-irrelevant distractors, but shift their reasoning by up to 6.5%, on average, towards the user's stated preferred moral view, and varying their reasoning depending on factors such as order (altering moral judgments by order in 13-22% of the cases) and duration (altering moral judgments between single-turn and multi-turn in 10-24% of the cases). Our analysis indicates that models tailor not just their final verdicts but their underlying justifications to align with a user's moral viewpoint - a failure mode we characterize as moral deliberative sycophancy.

📄 PDF Abstract BibTeX arXiv:2606.12731

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Logic-Parametric Neuro-Symbolic NLI: Controlling Logical Formalisms for Verifiable LLM Reasoning

2026-01-09 · Ali Farjami, Luca Redondi, Marco Valentino arxiv

Large language models (LLMs) and theorem provers (TPs) can be effectively combined for verifiable natural language inference (NLI). However, existing approaches rely on a fixed logical formalism, a feature that limits ro…

Natural Language Inference

Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare

2026-01-26 · Clément Christophe, Wadood Mohammed Abdul, Prateek Munjal, Tathagata Raha 외 arxiv

As LLMs are increasingly integrated into clinical workflows, their tendency for sycophancy, prioritizing user agreement over factual accuracy, poses significant risks to patient safety. While existing evaluations often r…

Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives

2025-10-30 · Kentaro Ozeki, Risako Ando, Takanobu Morishita, Hirohiko Abe 외 arxiv

Normative reasoning is a type of reasoning that involves normative or deontic modality, such as obligation and permission. While large language models (LLMs) have demonstrated remarkable performance across various reason…

Reasoning Core: A Scalable RL Environment for LLM Symbolic Reasoning

2025-09-22 · Valentin Lacombe, Valentin Quesnel, Damien Sileo arxiv

We introduce Reasoning Core, a new scalable environment for Reinforcement Learning with Verifiable Rewards (RLVR), designed to advance foundational symbolic reasoning in Large Language Models (LLMs). Unlike existing benc…

Reinforcement Learning

VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models

2026-06-15 · Sen Xu, Shixi Liu, Wei Wang, Jixin Min 외 arxiv

This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reasoning can be pushed within a strictly small-model regime. Building upon the Spectr…

Reinforcement Learning