paper-with-me

홈 › Papers

Demonstrations of Integrity Attacks in Multi-Agent Systems

2025-06-05 · Can Zheng, Yuhan Cao, Xiaoning Dong, Tianxing He

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding, code generation, and complex planning. Simultaneously, Multi-Agent Systems (MAS) have garnered attention for their potential to enable cooperation among distributed agents. However, from a multi-party perspective, MAS could be vulnerable to malicious agents that exploit the system to serve self-interests without disrupting its core functionality. This work explores integrity attacks where malicious agents employ subtle prompt manipulation to bias MAS operations and gain various benefits. Four types of attacks are examined: \textit{Scapegoater}, who misleads the system monitor to underestimate other agents' contributions; \textit{Boaster}, who misleads the system monitor to overestimate their own performance; \textit{Self-Dealer}, who manipulates other agents to adopt certain tools; and \textit{Free-Rider}, who hands off its own task to others. We demonstrate that strategically crafted prompts can introduce systematic biases in MAS behavior and executable instructions, enabling malicious agents to effectively mislead evaluation systems and manipulate collaborative agents. Furthermore, our attacks can bypass advanced LLM-based monitors, such as GPT-4o-mini and o3-mini, highlighting the limitations of current detection mechanisms. Our findings underscore the critical need for MAS architectures with robust security protocols and content validation mechanisms, alongside monitoring systems capable of comprehensive risk scenario assessment.

📄 PDF Abstract BibTeX arXiv:2506.04572

Code (0)

등록된 구현이 없습니다.

Tasks

Code GenerationNatural Language Understanding

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
ADOPT Please enter a description about the method here
MAS This optimizer mix ADAM and SGD creating the MAS optimizer.

Similar Papers 제목 키워드 기반

Compromising Embodied Agents with Contextual Backdoor Attacks

2024-08-06 · Aishan Liu, Yuguang Zhou, Xianglong Liu, Tianyuan Zhang 외

Large language models (LLMs) have transformed the development of embodied intelligence. By providing a few contextual demonstrations, developers can utilize the extensive internal knowledge of LLMs to effortlessly transl…

Autonomous DrivingRobot ManipulationVisual Reasoning

ACE: A Security Architecture for LLM-Integrated App Systems

2025-04-29 · Evan Li, Tushin Mallick, Evan Rose, William Robertson 외

LLM-integrated app systems extend the utility of Large Language Models (LLMs) with third-party apps that are invoked by a system LLM using interleaved planning and execution phases to answer user queries. These systems i…

Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems

2025-10-20 · Rishi Jha, Harold Triedman, Justin Wagle, Vitaly Shmatikov arxiv

Control-flow hijacking attacks manipulate orchestration mechanisms in multi-agent systems into performing unsafe actions that compromise the system and exfiltrate sensitive information. Recently proposed defenses, such a…

An information theoretic vulnerability metric for data integrity attacks on smart grids

2022-11-04 · Xiuzhen Ye, Iñaki Esnaola, Samir M. Perlaza, Robert F. Harrison

A novel metric that describes the vulnerability of the measurements in power systems to data integrity attacks is proposed. The new metric, coined vulnerability index (VuIx), leverages information theoretic measures to a…

Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems

2025-08-12 · Yutong Wu, Jie Zhang, Yiming Li, Chao Zhang 외 arxiv

Vision Language Model (VLM)-based agents are stateful, autonomous entities capable of perceiving and interacting with their environments through vision and language. Multi-agent systems comprise specialized agents who co…