paper-with-me

Papers

A Dynamic Stackelberg Game Framework for Agentic AI Defense Against LLM Jailbreaking

2025-07-10 · Zhengye Han, Quanyan Zhu

As large language models (LLMs) are increasingly deployed in critical applications, the challenge of jailbreaking, where adversaries manipulate the models to bypass safety mechanisms, has become a significant concern. This paper presents a dynamic Stackelberg game framework to model the interactions between attackers and defenders in the context of LLM jailbreaking. The framework treats the prompt-response dynamics as a sequential extensive-form game, where the defender, as the leader, commits to a strategy while anticipating the attacker's optimal responses. We propose a novel agentic AI solution, the "Purple Agent," which integrates adversarial exploration and defensive strategies using Rapidly-exploring Random Trees (RRT). The Purple Agent actively simulates potential attack trajectories and intervenes proactively to prevent harmful outputs. This approach offers a principled method for analyzing adversarial dynamics and provides a foundation for mitigating the risk of jailbreaking.

📄 PDF Abstract BibTeX arXiv:2507.08207

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Meta Stackelberg Game: Robust Federated Learning against Adaptive and Mixed Poisoning Attacks

2024-10-22 · Tao Li, Henger Li, Yunian Pan, Tianyi Xu 외

Federated learning (FL) is susceptible to a range of security threats. Although various defense mechanisms have been proposed, they are typically non-adaptive and tailored to specific types of attacks, leaving them insuf…

Federated LearningMeta-LearningModel PoisoningReinforcement Learning (RL)

LLM-Empowered Agentic MAC Protocols: A Dynamic Stackelberg Game Approach

2025-10-13 · Renxuan Tan, Rongpeng Li, Fei Wang, Chenghui Peng 외 arxiv

Medium Access Control (MAC) protocols, essential for wireless networks, are typically manually configured. While deep reinforcement learning (DRL)-based protocols enhance task-specified network performance, they suffer f…

Reinforcement Learning

Multi-agent Reinforcement Learning in Bayesian Stackelberg Markov Games for Adaptive Moving Target Defense

2020-07-20 · Sailik Sengupta, Subbarao Kambhampati

The field of cybersecurity has mostly been a cat-and-mouse game with the discovery of new attacks leading the way. To take away an attacker's advantage of reconnaissance, researchers have proposed proactive defense metho…

Multi-agent Reinforcement LearningQ-LearningReinforcement Learning (RL)

Convergence of Learning Dynamics in Stackelberg Games

2019-06-04 · Tanner Fiez, Benjamin Chasnov, Lillian J. Ratliff

This paper investigates the convergence of learning dynamics in Stackelberg games. In the class of games we consider, there is a hierarchical game being played between a leader and a follower with continuous action space…

Learning Movement Strategies for Moving Target Defense

2021-01-01 · Sailik Sengupta, Subbarao Kambhampati

The field of cybersecurity has mostly been a cat-and-mouse game with the discovery of new attacks leading the way. To take away an attacker's advantage of reconnaissance, researchers have proposed proactive defense metho…

Q-Learning