paper-with-me

홈 › Papers

Syntactic Framing Fragility: An Audit of Robustness in LLM Ethical Decisions

2025-12-27 · Katherine Elkins, Jon Chun arxiv

Large language models exhibit systematic negation sensitivity, yet no operational framework exists to measure this vulnerability at deployment scale, especially in high-stakes decisions. We introduce Syntactic Framing Fragility (SFF), a framework for quantifying decision consistency under logically equivalent syntactic transformations. SFF isolates syntactic effects via Logical Polarity Normalization, enabling direct comparison across positive and negative framings while controlling for polarity inversion, and provides the Syntactic Variation Index (SVI) as a robustness metric suitable for CI/CD integration. Auditing 23 models across 14 high-stakes scenarios (39,975 decisions), we establish ground-truth effect sizes for a phenomenon previously characterized only qualitatively and find that open-source models exhibit $2.2x higher fragility than commercial counterparts. Negation-bearing syntax is the dominant failure mode, with some models endorsing actions at 80-97% rates even when asked whether agents not act. These patterns are consistent with negation suppression failure documented in prior work, with chain-of-thought reasoning reducing fragility in some but not all cases. We provide scenario-stratified risk profiles and offer an operational checklist compatible with EU AI Act and NIST RMF requirements. Code, data, and scenarios will be released upon publication.

📄 PDF Abstract BibTeX arXiv:2601.09724

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

2026-01-29 · Katherine Elkins, Jon Chun arxiv

Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framing. We audit 16 models across 14 ethically fraught dilemmas using polarit…

Toward Constraint Compliant Goal Formulation and Planning

2024-05-21 · Steven J. Jones, Robert E. Wray

One part of complying with norms, rules, and preferences is incorporating constraints (such as knowledge of ethics) into one's goal formulation and planning processing. We explore in a simple domain how the encoding of k…

Ethics

The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models

2025-12-02 · Saeid Jamshidi, Kawser Wazed Nafi, Arghavan Moradi Dakhel, Negar Shahabi 외 arxiv

The rapid advancement and adaptability of Large Language Models (LLMs) highlight the need for moral consistency, the capacity to maintain ethically coherent reasoning across varied contexts. Existing alignment frameworks…

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs

2025-09-04 · Shei Pern Chua, Zhen Leng Thai, Kai Jun Teh, Xiao Li 외 arxiv

Large Language Model safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. This classification proves insufficient when models encounter ethical dilemmas, where the capac…

Towards An Ethics-Audit Bot

2021-03-29 · Siani Pearson, Martin Lloyd, Vivek Nallur

In this paper we focus on artificial intelligence (AI) for governance, not governance for AI, and on just one aspect of governance, namely ethics audit. Different kinds of ethical audit bots are possible, but who makes t…

Ethics