paper-with-me

홈 › Papers

In-Context Defense in Computer Agents: An Empirical Study

2025-03-12 · Pei Yang, Hai Ci, Mike Zheng Shou

Computer agents powered by vision-language models (VLMs) have significantly advanced human-computer interaction, enabling users to perform complex tasks through natural language instructions. However, these agents are vulnerable to context deception attacks, an emerging threat where adversaries embed misleading content into the agent's operational environment, such as a pop-up window containing deceptive instructions. Existing defenses, such as instructing agents to ignore deceptive elements, have proven largely ineffective. As the first systematic study on protecting computer agents, we introduce textbf{in-context defense}, leveraging in-context learning and chain-of-thought (CoT) reasoning to counter such attacks. Our approach involves augmenting the agent's context with a small set of carefully curated exemplars containing both malicious environments and corresponding defensive responses. These exemplars guide the agent to first perform explicit defensive reasoning before action planning, reducing susceptibility to deceptive attacks. Experiments demonstrate the effectiveness of our method, reducing attack success rates by 91.2% on pop-up window attacks, 74.6% on average on environment injection attacks, while achieving 100% successful defenses against distracting advertisements. Our findings highlight that (1) defensive reasoning must precede action planning for optimal performance, and (2) a minimal number of exemplars (fewer than three) is sufficient to induce an agent's defensive behavior.

📄 PDF Abstract BibTeX arXiv:2503.09241

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents

2025-06-03 · Tri Cao, Bennett Lim, Yue Liu, Yuan Sui 외

Computer-Use Agents (CUAs) with full system access enable powerful task automation but pose significant security and privacy risks due to their ability to manipulate files, access user data, and execute arbitrary command…

The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis

2026-02-11 · Peiran Wang, Xinfeng Li, Chong Xiang, Jinghuai Zhang 외 arxiv

The evolution of Large Language Models (LLMs) has resulted in a paradigm shift towards autonomous agents, necessitating robust security against Prompt Injection (PI) vulnerabilities where untrusted inputs hijack agent be…

Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents

2025-03-20 · Atharv Singh Patlan, Peiyao Sheng, S. Ashwin Hebbar, Prateek Mittal 외

The integration of AI agents with Web3 ecosystems harnesses their complementary potential for autonomy and openness yet also introduces underexplored security risks, as these agents dynamically interact with financial pr…

AI Agent

Automated Cyber Defense with Generalizable Graph-based Reinforcement Learning Agents

2025-09-19 · Isaiah J. King, Benjamin Bowman, H. Howie Huang arxiv

Deep reinforcement learning (RL) is emerging as a viable strategy for automated cyber defense (ACD). The traditional RL approach represents networks as a list of computers in various states of safety or threat. Unfortuna…

Reinforcement Learning

Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content

2026-05-23 · Rohan Pandey, Archit Bhujang arxiv

Large language models (LLMs) are increasingly used as analyst assistants in security operations centers (SOCs), where they ingest log and alert data to produce triage labels, incident summaries, or remediation advice. We…