paper-with-me

홈 › Papers

Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference

2024-06-21 · Anton Xue, Avishree Khare, Rajeev Alur, Surbhi Goel, Eric Wong

We study how to subvert large language models (LLMs) from following prompt-specified rules. We first formalize rule-following as inference in propositional Horn logic, a mathematical system in which rules have the form "if $P$ and $Q$, then $R$" for some propositions $P$, $Q$, and $R$. Next, we prove that although small transformers can faithfully follow such rules, maliciously crafted prompts can still mislead both theoretical constructions and models learned from data. Furthermore, we demonstrate that popular attack algorithms on LLMs find adversarial prompts and induce attention patterns that align with our theory. Our novel logic-based framework provides a foundation for studying LLMs in rule-based settings, enabling a formal analysis of tasks like logical reasoning and jailbreak attacks.

📄 PDF Abstract BibTeX arXiv:2407.00075

Code (0)

등록된 구현이 없습니다.

Tasks

Logical Reasoning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Subversion via Focal Points: Investigating Collusion in LLM Monitoring

2025-07-02 · Olli Järviniemi arxiv

We evaluate language models' ability to subvert monitoring protocols via collusion. More specifically, we have two instances of a model design prompts for a policy (P) and a monitor (M) in a programming task setting. The…

Governable AI: Provable Safety Under Extreme Threat Models

2025-08-28 · Donglin Wang, Weiyun Liang, Chunyuan Chen, Jing Xu 외 arxiv

As AI rapidly advances, the security risks posed by AI are becoming increasingly severe, especially in critical scenarios, including those posing existential risks. If AI becomes uncontrollable, manipulated, or actively …

Subversion Strategy Eval: Evaluating AI's stateless strategic capabilities against control protocols

2024-12-17 · Alex Mallen, Charlie Griffin, Alessandro Abate, Buck Shlegeris

AI control protocols are plans for usefully deploying AI systems in a way that is safe, even if the AI intends to subvert the protocol. Previous work evaluated protocols by subverting them with a human-AI red team, where…

Subversive Characters and Stereotyping Readers: Characterizing Queer Relationalities with Dialogue-Based Relation Extraction

2024-10-19 · Kent K. Chang, Anna Ho, David Bamman

Television is often seen as a site for subcultural identification and subversive fantasy, including in queer cultures. How might we measure subversion, or the degree to which the depiction of social relationship between …

RelationRelation Extraction

Gradient-based Data Subversion Attack Against Binary Classifiers

2021-05-31 · Rosni K Vasu, Sanjay Seetharaman, Shubham Malaviya, Manish Shukla 외

Machine learning based data-driven technologies have shown impressive performances in a variety of application domains. Most enterprises use data from multiple sources to provide quality applications. The reliability of …

BIG-bench Machine LearningData Poisoning