paper-with-me

홈 › Papers

HARP: Measuring Harm Amplification in Multi-Agent LLM Systems

2026-05-26 · Md Hafizur Rahman, Zafaryab Haider, Tanzim Mahfuz, Prabuddha Chakraborty arxiv

Multi-agent LLM systems decompose workflows across agents, tools, shared context, memory, and decision gates. This modularity improves interpretability, but creates a propagation risk: a bounded perturbation to one component can be reused by other agents and amplified into system-level harm. We introduce HARP (Harm Amplification through Role Perturbation), a trace-first methodology for studying local-to-global harm amplification in multi-agent LLM systems. HARP compares paired clean and perturbed executions and records specialist outputs, tool calls, memory reads/writes, guard events, oracle logs, latency, token cost, and decisions. We define local harm as deviation from targeted agents or corrupted channels, global harm as deviation over the full trace, and harm amplification as (H_global/H_local). This complements attack success rate with a measure of how strongly orchestration spreads harm beyond the attack point. We instantiate HARP in a finance-oriented seven-agent system with a deterministic decision gate and configurable attack harness for specialist compromise, collusion, shared-context corruption, and temporal or memory-persistent attacks. Across five defenses, prompt-only defenses preserve benign utility but leave high success and stealth; pre-tool and step-level guards reduce some failures with utility or latency costs; and IntegrityGuard, a trace-consistency defense, achieves the lowest attack success and global harm but introduces utility/cost trade-offs. Results show that single-specialist compromise produces the strongest amplification, shared-context corruption yields the highest attack success, and temporal persistence produces the largest malicious impact. HARP argues that secure multi-agent evaluation must measure not only bypass, but propagation.

📄 PDF Abstract BibTeX arXiv:2605.27489

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

2024-10-11 · Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas 외

The robustness of LLMs to jailbreak attacks, where users design prompts to circumvent safety measures and misuse model capabilities, has been studied primarily for LLMs acting as simple chatbots. Meanwhile, LLM agents --…

SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models

2026-01-29 · Alok Abhishek, Tushar Bandopadhyay, Lisa Erickson arxiv

Large language models (LLMs) are increasingly deployed in high-stakes domains, where rare but severe failures can result in irreversible harm. However, prevailing evaluation benchmarks often reduce complex social risk to…

Investigating and Alleviating Harm Amplification in LLM Interactions

2026-06-01 · Ruohao Guo, Wei Xu, Alan Ritter arxiv

Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve harmful outcomes beyond their capabilities through extended interactio…

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

2026-07-08 · Lifei Liu, Haoran Yu, Xiaochong Jiang, Su Wang 외 arxiv

Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret…

Harm Amplification in Text-to-Image Models

2024-02-01 · Susan Hao, Renee Shelby, Yuchi Liu, Hansa Srinivasan 외

Text-to-image (T2I) models have emerged as a significant advancement in generative AI; however, there exist safety concerns regarding their potential to produce harmful image outputs even when users input seemingly safe …