paper-with-me

Papers

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

2026-04-11 · Vishal Pramanik, Maisha Maliha, Susmit Jha, Sumit Kumar Jha arxiv

Large language models remain vulnerable to jailbreak attacks -- inputs designed to bypass safety mechanisms and elicit harmful responses -- despite advances in alignment and instruction tuning. We propose Head-Masked Nullspace Steering (HMNS), a circuit-level intervention that (i) identifies attention heads most causally responsible for a model's default behavior, (ii) suppresses their write paths via targeted column masking, and (iii) injects a perturbation constrained to the orthogonal complement of the muted subspace. HMNS operates in a closed-loop detection-intervention cycle, re-identifying causal heads and reapplying interventions across multiple decoding attempts. Across multiple jailbreak benchmarks, strong safety defenses, and widely used language models, HMNS attains state-of-the-art attack success rates with fewer queries than prior methods. Ablations confirm that nullspace-constrained injection, residual norm scaling, and iterative re-identification are key to its effectiveness. To our knowledge, this is the first jailbreak method to leverage geometry-aware, interpretability-informed interventions, highlighting a new paradigm for controlled model steering and adversarial safety circumvention.

📄 PDF Abstract BibTeX arXiv:2604.10326

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Probe-based Steering for Robust LLM Jailbreaking

2026-05-19 · Junxi Chen, Junhao Dong, Xiaohua Xie arxiv

Recent work has demonstrated the potential of contrastive steering for jailbreaking Large Language Models (LLMs). However, existing methods rely on limited and inherently biased contrastive prompts and require laborious …

Model extraction

Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP

2026-06-11 · Elisabetta Rocchetti, Alfio Ferrara arxiv

Arditi et al. (2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual stream, recoverable by a difference-in-means (DiM) of harmful and harmless activations…

Multiple Correlated Jammers Nullification using LSTM-based Deep Dueling Neural Network

2022-02-08 · Linh Manh Hoang, Diep N. Nguyen, J. Andrew Zhang, Dinh Thai Hoang

Suppressing the deliberate interference for wireless networks is critical to guarantee a reliable communication link. However, nullifying the jamming signals can be problematic when the correlations between transmitted j…

Q-Learning

NullSpaceNet: Nullspace Convoluional Neural Network with Differentiable Loss Function

2020-04-25 · Mohamed H. Abdelpakey, Mohamed S. Shehata

We propose NullSpaceNet, a novel network that maps from the pixel level input to a joint-nullspace (as opposed to the traditional feature space), where the newly learned joint-nullspace features have clearer interpretati…

Fixed Horizon Linear Quadratic Covariance Steering in Continuous Time with Hilbert-Schmidt Terminal Cost

2025-10-24 · Tushar Sial, Abhishek Halder arxiv

We formulate and solve the fixed horizon linear quadratic covariance steering problem in continuous time with a terminal cost measured in Hilbert-Schmidt (i.e., Frobenius) norm error between the desired and the controlle…