paper-with-me

홈 › Papers

Matching Ranks Over Probability Yields Truly Deep Safety Alignment

2025-12-05 · Jason Vega, Gagandeep Singh arxiv

A frustratingly easy technique known as the prefilling attack has been shown to effectively circumvent the safety alignment of frontier LLMs by simply prefilling the assistant response with an affirmative prefix before decoding. In response, recent work proposed a supervised fine-tuning (SFT) defense using data augmentation to achieve a \enquote{deep} safety alignment, allowing the model to generate natural language refusals immediately following harmful prefills. Unfortunately, we show in this work that the "deep" safety alignment produced by such an approach is in fact not very deep. A generalization of the prefilling attack, which we refer to as the Rank-Assisted Prefilling (RAP) attack, can effectively extract harmful content from models fine-tuned with the data augmentation defense by selecting low-probability "harmful" tokens from the top 20 predicted next tokens at each step (thus ignoring high-probability "refusal" tokens). We argue that this vulnerability is enabled due to the "gaming" of the SFT objective when the target distribution entropies are low, where low fine-tuning loss is achieved by shifting large probability mass to a small number of refusal tokens while neglecting the high ranks of harmful tokens. We then propose a new perspective on achieving deep safety alignment by matching the token ranks of the target distribution, rather than their probabilities. This perspective yields a surprisingly simple fix to the data augmentation defense based on regularizing the attention placed on harmful prefill tokens, an approach we call PRefill attEntion STOpping (PRESTO). Adding PRESTO yields up to a 4.7x improvement in the mean StrongREJECT score under RAP attacks across three popular open-source LLMs, with low impact to model utility.

📄 PDF Abstract BibTeX arXiv:2512.05518

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

Distributional conformal prediction

2019-09-17 · Victor Chernozhukov, Kaspar Wüthrich, Yinchu Zhu

We propose a robust method for constructing conditionally valid prediction intervals based on models for conditional distributions such as quantile and distribution regression. Our approach can be applied to important pr…

Conformal PredictioncounterfactualPredictionPrediction Intervals+4

MAP Disparity Estimation Using Hidden Markov Trees

2015-12-01 · ICCV 2015 12 · Eric T. Psota, Jedrzej Kowalczuk, Mateusz Mittek, Lance C. Perez

A new method is introduced for stereo matching that operates on minimum spanning trees (MSTs) generated from the images. Disparity maps are represented as a collection of hidden states on MSTs, and each MST is modeled as…

Disparity EstimationStereo MatchingStereo Matching Hand

Generalized Wasserstein Flow Matching: Transport Plans, Everywhere, All at Once

2026-05-08 · Moritz Piening, Richard Duong, Gabriele Steidl arxiv

Flow matching has recently emerged as a flexible and efficient framework for generative modelling by learning deterministic transport dynamics between probability measures. In this work, we extend flow matching to the sp…

On the Identifiability of Tensor Ranks via Prior Predictive Matching

2025-10-16 · Eliezer da Silva, Arto Klami, Diego Mesquita, Iñigo Urteaga arxiv

Selecting the latent dimensions (ranks) in tensor factorization is a central challenge that often relies on heuristic methods. This paper introduces a rigorous approach to determine rank identifiability in probabilistic …

Focus of Attention Improves Information Transfer in Visual Features

2020-06-16 · NeurIPS 2020 12 · Matteo Tiezzi, Stefano Melacci, Alessandro Betti, Marco Maggini 외

Unsupervised learning from continuous visual streams is a challenging problem that cannot be naturally and efficiently managed in the classic batch-mode setting of computation. The information stream must be carefully pr…