paper-with-me

Papers

Alignment Under Pressure: The Case for Informed Adversaries When Evaluating LLM Defenses

2025-05-21 · Xiaoxue Yang, Bozhidar Stevanoski, Matthieu Meeus, Yves-Alexandre de Montjoye

Large language models (LLMs) are rapidly deployed in real-world applications ranging from chatbots to agentic systems. Alignment is one of the main approaches used to defend against attacks such as prompt injection and jailbreaks. Recent defenses report near-zero Attack Success Rates (ASR) even against Greedy Coordinate Gradient (GCG), a white-box attack that generates adversarial suffixes to induce attacker-desired outputs. However, this search space over discrete tokens is extremely large, making the task of finding successful attacks difficult. GCG has, for instance, been shown to converge to local minima, making it sensitive to initialization choices. In this paper, we assess the future-proof robustness of these defenses using a more informed threat model: attackers who have access to some information about the alignment process. Specifically, we propose an informed white-box attack leveraging the intermediate model checkpoints to initialize GCG, with each checkpoint acting as a stepping stone for the next one. We show this approach to be highly effective across state-of-the-art (SOTA) defenses and models. We further show our informed initialization to outperform other initialization methods and show a gradient-informed checkpoint selection strategy to greatly improve attack performance and efficiency. Importantly, we also show our method to successfully find universal adversarial suffixes -- single suffixes effective across diverse inputs. Our results show that, contrary to previous beliefs, effective adversarial suffixes do exist against SOTA alignment-based defenses, that these can be found by existing attack methods when adversaries exploit alignment knowledge, and that even universal suffixes exist. Taken together, our results highlight the brittleness of current alignment-based methods and the need to consider stronger threat models when testing the safety of LLMs.

📄 PDF Abstract BibTeX arXiv:2505.15738

Code (1)

computationalprivacy/checkpoint-gcg 공식 구현 jax

Similar Papers 제목 키워드 기반

Physics-informed reservoir characterization from bulk and extreme pressure events with a differentiable simulator

2026-04-14 · Harun Ur Rashid, Mingxin Li, Aleksandra Pachalieva, Georg Stadler 외 arxiv

Accurate characterization of subsurface heterogeneity is challenging but essential for applications such as reservoir pressure management, geothermal energy extraction and CO$_2$, H$_2$, and wastewater injection operatio…

Physics-informed neural networks for solving Reynolds-averaged Navier-Stokes equations

2021-07-22 · Hamidreza Eivazi, Mojtaba Tahani, Philipp Schlatter, Ricardo Vinuesa

Physics-informed neural networks (PINNs) are successful machine-learning methods for the solution and identification of partial differential equations (PDEs). We employ PINNs for solving the Reynolds-averaged Navier-Stok…

The AI Alignment Paradox

2024-05-31 · Robert West, Roland Aydin

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today'…

BP-DeepONet: A new method for cuffless blood pressure estimation using the physcis-informed DeepONet

2024-02-29 · Lingfeng li, Xue-Cheng Tai, Raymond Chan

Cardiovascular diseases (CVDs) are the leading cause of death worldwide, with blood pressure serving as a crucial indicator. Arterial blood pressure (ABP) waveforms provide continuous pressure measurements throughout the…

Blood pressure estimationDiagnosticMeta-Learning

Correlated-informed neural networks: a new machine learning framework to predict pressure drop in micro-channels

2022-01-19 · J. A. Montanez-Barrera, J. M. Barroso-Maldonado, A. F. Bedoya-Santacruz, Adrian Mota-Babiloni

Accurate pressure drop estimation in forced boiling phenomena is important during the thermal analysis and the geometric design of cryogenic heat exchangers. However, current methods to predict the pressure drop have one…

Transfer Learning