paper-with-me

Papers

Minimizing Collateral Damage in Activation Steering

2026-05-01 · Tam Nguyen, Tu Anh Nguyen, Sina Alemohammad, Richard G. Baraniuk arxiv

Activation steering is a method for controlling Large Language Model (LLM) behavior by intervening in its internal representations to increase the alignment with a specific target feature direction. However, standard interventions, such as vector addition, often cause ``collateral damage", defined as unintended changes in the alignment of activations along other non-target feature directions. This damage occurs because standard methods implicitly assume the isotropy of non-target features. In this work, we provide a mathematical formalization of collateral damage and introduce a principled framework that models steering as a constrained optimization problem. Our method finds a new activation that minimizes the expected squared collateral change weighted by the empirical second-moment matrix of activations. This weighting encodes the nonuniform cost of the perturbation in different feature directions, in contrast to isotropic approaches that penalize changes uniformly in all feature directions. By accounting for the empirical second-moment of activations, our approach achieves more precise control while reducing the degradation of model performance on unrelated tasks.

📄 PDF Abstract BibTeX arXiv:2605.01167

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SwordBench: Evaluating Orthogonality of Steering Image Representations

2026-05-10 · Vladimir Zaigrajew, Dawid Pludowski, Hubert Baniecki, Przemyslaw Biecek arxiv

Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existing evaluation protocols are limited to ambiguous language modeling task…

Surgical Refusal Ablation: Disentangling Safety from Intelligence via Concept-Guided Spectral Cleaning

2026-01-13 · Tony Cristofano arxiv

Safety-aligned language models systematically refuse harmful requests. While activation steering can modulate refusal, ablating the raw "refusal vector" calculated from contrastive harmful and harmless prompts often caus…

Pre-Intervention Prediction of Sparse Autoencoder Steering Side Effects

2026-06-06 · Evan Duan arxiv

Sparse autoencoder (SAE) features are increasingly used to steer language models, but feature steering is rarely clean: the same intervention can behave inconsistently across contexts and perturb unrelated features. We i…

PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning

2026-06-16 · Bo Su, Ankit Shah, Thai Le arxiv

Machine unlearning for large language models (LLMs) aims to remove specified knowledge while preserving the rest of the model's capabilities. However, the boundary between knowledge to forget and knowledge to retain is o…

Multi-property Steering of Large Language Models with Dynamic Activation Composition

2024-06-25 · Daniel Scalena, Gabriele Sarti, Malvina Nissim

Activation steering methods were shown to be effective in conditioning language model generation by additively intervening over models' intermediate representations. However, the evaluation of these techniques has so far…

Language ModelingLanguage Modelling