paper-with-me

홈 › Papers

SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning

2026-06-01 · Lichao Wang, Zhaoxing Ren, Tianzhuo Yang, Jiaming Ji, Chi Harold Liu, Yaodong Yang, Juntao Dai arxiv

As Large Language Model (LLM) agents increasingly leverage the Model Context Protocol (MCP) to operate in complex environments, the expansion of their action spaces offers agents unsafe capabilities and underscores the risk of power-seeking. While broad action space and greater environment influence are essential for task fulfillment, they create a fragile risk surface where minor errors or hallucinations are magnified into catastrophic failures. In response, we propose SafeMCP, a {server-side} defense plugin that constrains tool acquisition via predictive reasoning regarding future safety risks. SafeMCP utilizes an internal world model for look-ahead reasoning to implement a two-tier defense: proactive tool filtering to constrain hazardous power expansion and immediate intervention as a fail-safe. To train SafeMCP, we introduce a three-stage pipeline comprising environmental dynamic grounding, safe policy initialization, and reinforcement learning (RL) with dual verifiable rewards. Experiments on PowerSeeking Bench, ToolEmu, and AgentHarm show that SafeMCP achieves a safe equilibrium, effectively mitigating risks while preserving agent utility.

📄 PDF Abstract BibTeX arXiv:2606.01991

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems

2025-06-16 · Junfeng Fang, Zijun Yao, Ruipeng Wang, Haokai Ma 외

The development of large language models (LLMs) has entered in a experience-driven era, flagged by the emergence of environment feedback-driven learning via reinforcement learning and tool-using agents. This encourages t…

PositionRed Teaming

SAIGuard: Communication-State Simulation for Proactive Defense of LLM Multi-Agent Systems

2026-06-10 · Ruxue Shi, Yili Wang, Mengnan Du, Qinggang Zhang 외 arxiv

LLM-based multi-agent systems (MAS) solve complex tasks through inter-agent collaboration, but their communication-driven nature also allows security risks to spread across agents and trigger system-wide failures. Existi…

Contextualized Privacy Defense for LLM Agents

2026-03-03 · Yule Wen, Yanzhe Zhang, Jianxun Lian, Xiaoyuan Yi 외 arxiv

LLM agents increasingly act on users' personal information, yet existing privacy defenses remain limited in both design and adaptability. Most prior approaches rely on static or passive defenses, such as prompting and gu…

Reinforcement Learning

Designing Ethical Learning for Agentic AI: Toegye Yi Hwang's Ethical Emotion Regulation Framework

2026-04-07 · Ji Yeon Kim arxiv

Agentic AI systems capable of autonomous goal setting and proactive intervention introduce new challenges for regulating moral-emotional processes in learning environments. Existing frameworks typically treat emotion as …

Toward Intelligent and Secure Cloud: Large Language Model Empowered Proactive Defense

2024-12-30 · Yuyang Zhou, Guang Cheng, Kang Du, Zihan Chen 외

The rapid evolution of cloud computing technologies and the increasing number of cloud applications have provided a large number of benefits in daily lives. However, the diversity and complexity of different components p…

Cloud ComputingCode GenerationDiversityLanguage Modeling+2