paper-with-me

홈 › Papers

Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM

2025-11-24 · Adarsh Kumarappan, Ayushi Mehrotra arxiv

The SmoothLLM defense provides a certification guarantee against jailbreaking attacks, but it relies on a strict "k-unstable" assumption that rarely holds in practice. This strong assumption can limit the trustworthiness of the provided safety certificate. In this work, we address this limitation by introducing a more realistic probabilistic framework, "(k, $\varepsilon$)-unstable," to certify defenses against diverse jailbreaking attacks, from gradient-based (GCG) to semantic (PAIR). We derive a new, data-informed lower bound on SmoothLLM's defense probability by incorporating empirical models of attack success, providing a more trustworthy and practical safety certificate. By introducing the notion of (k, $\varepsilon$)-unstable, our framework provides practitioners with actionable safety guarantees, enabling them to set certification thresholds that better reflect the real-world behavior of LLMs. Ultimately, this work contributes a practical and theoretically-grounded mechanism to make LLMs more resistant to the exploitation of their safety alignments, a critical challenge in secure AI deployment.

📄 PDF Abstract BibTeX arXiv:2511.18721

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VeRecycle: Reclaiming Guarantees from Probabilistic Certificates for Stochastic Dynamical Systems after Change

2025-05-20 · Sterre Lutz, Matthijs T. J. Spaan, Anna Lukina

Autonomous systems operating in the real world encounter a range of uncertainties. Probabilistic neural Lyapunov certification is a powerful approach to proving safety of nonlinear stochastic dynamical systems. When face…

Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

2026-06-03 · Mohit Prashant, Arvind Easwaran arxiv

Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep RL may demonstrate susceptibility to transition perturbations that resu…

Reinforcement Learning

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

2023-10-05 · Alexander Robey, Eric Wong, Hamed Hassani, George J. Pappas

Despite efforts to align large language models (LLMs) with human intentions, widely-used LLMs such as GPT, Llama, and Claude are susceptible to jailbreaking attacks, wherein an adversary fools a targeted LLM into generat…

Continuous-time Data-driven Barrier Certificate Synthesis

2025-03-17 · Luke Rickard, Alessandro Abate, Kostas Margellos

We consider the problem of verifying safety for continuous-time dynamical systems. Developing upon recent advancements in data-driven verification, we use only a finite number of sampled trajectories to learn a barrier c…

Data-Driven Neural Certificate Synthesis

2025-02-08 · Luke Rickard, Alessandro Abate, Kostas Margellos

We investigate the problem of verifying different properties of discrete time dynamical systems, namely, reachability, safety and reach-while-avoid. To achieve this, we adopt a data driven perspective and using past syst…

Generalization BoundsPAC learning