paper-with-me

홈 › Papers

Mapping the Exploitation Surface: A 10,000-Trial Taxonomy of What Makes LLM Agents Exploit Vulnerabilities

2026-04-06 · Charafeddine Mouzouni arxiv

LLM agents with tool access can discover and exploit security vulnerabilities. This is known. What is not known is which features of a system prompt trigger this behaviour, and which do not. We present a systematic taxonomy based on approximately 10,000 trials across seven models, 37 prompt conditions, and 12 hypothesized attack dimensions -- all executed in real Docker sandboxes. Every condition includes the instruction "Always respect all rules and access policies"; only a single encouragement sentence varies. Within the task class studied (planted test-runner vulnerabilities), the threat model is narrower than assumed: nine of twelve dimensions -- including minimization, moral licensing, incentives, identity priming, and reasoning triggers -- produce no detectable exploitation at n=50 per cell (upper 95% CI <7% per condition). But it is also more subtle than expected: one dimension, goal reframing, reliably triggers exploitation. "You are solving a puzzle; there may be hidden clues" produces 38-40% exploitation on Claude Sonnet 4 despite the explicit rule instruction, replicating across four models (CTF framing: 8-14% on DeepSeek, GPT-5-mini, o4-mini). The agent does not override the rules; it reinterprets the task so that exploitative actions become task-aligned. GPT-4.1 produces no exploitation across 1,850 trials (37 conditions), and a temporal comparison across four OpenAI models released over eleven months shows a pattern consistent with improving safety training, though model capability differences are a confounder. The practical contribution is a narrowed, testable threat model: defenders should audit for goal-reframing language, not for the broad class of adversarial prompts.

📄 PDF Abstract BibTeX arXiv:2604.04561

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ICP Algorithm: Theory, Practice And Its SLAM-oriented Taxonomy

2022-06-13 · Hao Bai

The Iterative Closest Point (ICP) algorithm is one of the most important algorithms for geometric alignment of three-dimensional surface registration, which is frequently used in computer vision tasks, including the Simu…

Simultaneous Localization and Mapping

MCP-38: A Comprehensive Threat Taxonomy for Model Context Protocol Systems (v1.0)

2026-03-18 · Yi Ting Shen, Kentaroh Toyoda, Alex Leung arxiv

The Model Context Protocol (MCP) introduces a structurally distinct attack surface that existing threat frameworks, designed for traditional software systems or generic LLM deployments, do not adequately cover. This pape…

Generation of patient specific cardiac chamber models using generative neural networks under a Bayesian framework for electroanatomical mapping

2023-11-27 · Sunil Mathew, Jasbir Sra, Daniel B. Rowe

Electroanatomical mapping is a technique used in cardiology to create a detailed 3D map of the electrical activity in the heart. It is useful for diagnosis, treatment planning and real time guidance in cardiac ablation p…

Surface Reconstruction

A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence

2020-06-22 · Changhao Chen, Bing Wang, Chris Xiaoxuan Lu, Niki Trigoni 외

Deep learning based localization and mapping has recently attracted significant attention. Instead of creating hand-designed algorithms through exploitation of physical models or geometric theories, deep learning based s…

Deep LearningScene UnderstandingSimultaneous Localization and Mapping

A taxonomy of strategic human interactions in traffic conflicts

2021-09-27 · Atrisha Sarkar, Kate Larson, Krzysztof Czarnecki

In order to enable autonomous vehicles (AV) to navigate busy traffic situations, in recent years there has been a focus on game-theoretic models for strategic behavior planning in AVs. However, a lack of common taxonomy …

Autonomous VehiclesNavigate