paper-with-me

Papers

Do LLMs Follow Their Own Rules? A Reflexive Audit of Self-Stated Safety Policies

2026-04-10 · Avni Mittal arxiv

LLMs internalize safety policies through RLHF, yet these policies are never formally specified and remain difficult to inspect. Existing benchmarks evaluate models against external standards but do not measure whether models understand and enforce their own stated boundaries. We introduce the Symbolic-Neural Consistency Audit (SNCA), a framework that (1) extracts a model's self-stated safety rules via structured prompts, (2) formalizes them as typed predicates (Absolute, Conditional, Adaptive), and (3) measures behavioral compliance via deterministic comparison against harm benchmarks. Evaluating four frontier models across 45 harm categories and 47,496 observations reveals systematic gaps between stated policy and observed behavior: models claiming absolute refusal frequently comply with harmful prompts, reasoning models achieve the highest self-consistency but fail to articulate policies for 29% of categories, and cross-model agreement on rule types is remarkably low (11%). These results demonstrate that the gap between what LLMs say and what they do is measurable and architecture-dependent, motivating reflexive consistency audits as a complement to behavioral benchmarks.

📄 PDF Abstract BibTeX arXiv:2604.09189

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AuditGPT: Auditing Smart Contracts with ChatGPT

2024-04-05 · Shihao Xia, Shuai Shao, Mengting He, Tingting Yu 외

To govern smart contracts running on Ethereum, multiple Ethereum Request for Comment (ERC) standards have been developed, each containing a set of rules to guide the behaviors of smart contracts. Violating the ERC rules …

SymGPT: Auditing Smart Contracts via Combining Symbolic Execution with Large Language Models

2025-02-11 · Shihao Xia, Mengting He, Shuai Shao, Tingting Yu 외

To govern smart contracts running on Ethereum, multiple Ethereum Request for Comment (ERC) standards have been developed, each having a set of rules to guide the behaviors of smart contracts. Violating the ERC rules coul…

Natural Language Understanding

Towards a Semi-Automatic Detection of Reflexive and Reciprocal Constructions and Their Representation in a Valency Lexicon

2020-05-01 · LREC 2020 5 · V{\'a}clava Kettnerov{\'a}, Marketa Lopatkova, Anna Vernerov{\'a}, Petra Barancikova

Valency lexicons usually describe valency behavior of verbs in non-reflexive and non-reciprocal constructions. However, reflexive and reciprocal constructions are common morphosyntactic forms of verbs. Both of these cons…

Word Embeddings

GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents

2025-05-16 · Lingxiao Diao, Xinyue Xu, Wanxuan Sun, Cheng Yang 외

Large language models (LLMs) have been widely deployed as autonomous agents capable of following user instructions and making decisions in real-world applications. Previous studies have made notable progress in benchmark…

BenchmarkingInstruction Following

An Efficient Edge Detection Technique by Two Dimensional Rectangular Cellular Automata

2013-12-22 · Jahangir Mohammed, Deepak Ranjan Nayak

This paper proposes a new pattern of two dimensional cellular automata linear rules that are used for efficient edge detection of an image. Since cellular automata is inherently parallel in nature, it has produced desire…

Edge DetectionVocal Bursts Valence Prediction