paper-with-me

홈 › Papers

ToolTweak: An Attack on Tool Selection in LLM-based Agents

2025-10-02 · Jonathan Sneh, Ruomei Yan, Jialin Yu, Philip Torr, Yarin Gal, Sunando Sengupta, Eric Sommerlade, Alasdair Paren, Adel Bibi arxiv

As LLMs increasingly power agents that interact with external tools, tool use has become an essential mechanism for extending their capabilities. These agents typically select tools from growing databases or marketplaces to solve user tasks, creating implicit competition among tool providers and developers for visibility and usage. In this paper, we show that this selection process harbors a critical vulnerability: by iteratively manipulating tool names and descriptions, adversaries can systematically bias agents toward selecting specific tools, gaining unfair advantage over equally capable alternatives. We present ToolTweak, a lightweight automatic attack that increases selection rates from a baseline of around 20% to as high as 81%, with strong transferability between open-source and closed-source models. Beyond individual tools, we show that such attacks cause distributional shifts in tool usage, revealing risks to fairness, competition, and security in emerging tool ecosystems. To mitigate these risks, we evaluate two defenses: paraphrasing and perplexity filtering, which reduce bias and lead agents to select functionally similar tools more equally. All code will be open-sourced upon acceptance.

📄 PDF Abstract BibTeX arXiv:2510.02554

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning

2026-05-24 · Xuanye Zhang, Yongsen Zheng, Zhuqin Xu, Kaiyu Zhou 외 arxiv

LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents toward inappropriate/wrong tools and enabling malicious actions. Most …

ToolFlood: Beyond Selection -- Hiding Valid Tools from LLM Agents via Semantic Covering

2026-03-14 · Hussein Jawad, Nicolas J-B Brunel arxiv

Large Language Model (LLM) agents increasingly use external tools for complex tasks and rely on embedding-based retrieval to select a small top-k subset for reasoning. As these systems scale, the robustness of this retri…

Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious Tools

2025-08-04 · Kanghua Mo, Li Hu, Yucheng Long, Zhihao Li arxiv

Large language model (LLM) agents have demonstrated remarkable capabilities in complex reasoning and decision-making by leveraging external tools. However, this tool-centric paradigm introduces a previously underexplored…

InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

2024-03-05 · Qiusi Zhan, Zhixiang Liang, Zifan Ying, Daniel Kang

Recent work has embodied LLMs as agents, allowing them to access tools, perform actions, and interact with external content (e.g., emails or websites). However, external content introduces the risk of indirect prompt inj…

BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model

MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents

2025-10-14 · Dongsen Zhang, Zekun Li, Xu Luo, Xuannan Liu 외 arxiv

The Model Context Protocol (MCP) standardizes how large language model (LLM) agents discover, describe, and call external tools. While MCP unlocks broad interoperability, it also enlarges the attack surface by making too…

Instruction Following