paper-with-me

홈 › Papers

Optimizing AI Agent Attacks With Synthetic Data

2025-11-04 · Chloe Loughridge, Paul Colognese, Avery Griffin, Tyler Tracy, Jon Kutasov, Joe Benton arxiv

As AI deployments become more complex and high-stakes, it becomes increasingly important to be able to estimate their risk. AI control is one framework for doing so. However, good control evaluations require eliciting strong attack policies. This can be challenging in complex agentic environments where compute constraints leave us data-poor. In this work, we show how to optimize attack policies in SHADE-Arena, a dataset of diverse realistic control environments. We do this by decomposing attack capability into five constituent skills -- suspicion modeling, attack selection, plan synthesis, execution and subtlety -- and optimizing each component individually. To get around the constraint of limited data, we develop a probabilistic model of attack dynamics, optimize our attack hyperparameters using this simulation, and then show that the results transfer to SHADE-Arena. This results in a substantial improvement in attack strength, reducing safety score from a baseline of 0.87 to 0.41 using our scaffold.

📄 PDF Abstract BibTeX arXiv:2511.02823

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks

2024-12-03 · Zijiao Yang, Xiangxi Shi, Eric Slyman, Stefan Lee

Assistive embodied agents that can be instructed in natural language to perform tasks in open-world environments have the potential to significantly impact labor tasks like manufacturing or in-home care -- benefiting the…

Adversarial AttackVision and Language Navigation

Byzantine-Resilient Output Optimization of Multiagent via Self-Triggered Hybrid Detection Approach

2024-10-17 · Chenhang Yan, Liping Yan, Yuezu Lv, Bolei Dong 외

How to achieve precise distributed optimization despite unknown attacks, especially the Byzantine attacks, is one of the critical challenges for multiagent systems. This paper addresses a distributed resilient optimizati…

Distributed Optimizationvalid

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors

2026-07-08 · Oliver Makins, Orazio Angelini, Zohreh Shams, Mary Phuong arxiv

AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent. AI Control usually studies a single agent in one trajectory, but real deployments run many agents over sha…

Autodata: An agentic data scientist to create high quality synthetic data

2026-06-24 · Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie 외 arxiv

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it l…

Legal Reasoning

Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution

2025-06-01 · Meysam Alizadeh, Zeynab Samei, Daria Stetsenko, Fabrizio Gilardi

Previous benchmarks on prompt injection in large language models (LLMs) have primarily focused on generic tasks and attacks, offering limited insights into more complex threats like data exfiltration. This paper examines…