paper-with-me

홈 › Papers

Predicting Where Steering Vectors Succeed

2026-04-16 · Jayadev Billa arxiv

Steering vectors work for some concepts and layers but fail for others, and practitioners have no way to predict which setting applies before running an intervention. We introduce the Linear Accessibility Profile (LAP), a per-layer diagnostic that repurposes the logit lens as a predictor of steering vector effectiveness. The key measure, $A_{\mathrm{lin}}$, applies the model's unembedding matrix to intermediate hidden states, requiring no training. Across 24 controlled binary concept families on five models (Pythia-2.8B to Llama-8B), peak $A_{\mathrm{lin}}$ predicts steering effectiveness at $ρ= +0.86$ to $+0.91$ and layer selection at $ρ= +0.63$ to $+0.92$. A three-regime framework explains when difference-of-means steering works, when nonlinear methods are needed, and when no method can work. An entity-steering demo confirms the prediction end-to-end: steering at the LAP-recommended layer redirects completions on Gemma-2-2B and OLMo-2-1B-Instruct, while the middle layer (the standard heuristic) has no effect on either model.

📄 PDF Abstract BibTeX arXiv:2604.15557

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

2026-04-03 · Mohammed Suhail B Nadaf arxiv

Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both steer the model and be readable along the unembedding. Function vectors (FVs…

Cognitive Steering in Deep Neural Networks via Long-Range Modulatory Feedback Connections

2023-09-21 · NeurIPS 2023 11

Given the rich visual information available in each glance, humans can internally direct their visual attention to enhance goal-relevant information---a capacity often absent in standard vision models. Here we introduce…

Adversarial Robustness of Activation Steering in Large Language Models

2026-06-05 · Kien Le, Thai Le arxiv

Activation steering has become a popular training-free method to control LLM behavior by injecting precomputed direction vectors into the model's residual stream at inference time. Yet its robustness to realistic input v…

Adversarial Robustness

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

2026-04-09 · Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha arxiv

Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explanation for how it works-- specifically, what internal mechanisms steeri…

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

2026-07-22 · Kavin Aravindan, Arihant Rastogi, Aadi Prasad, Krishak Aneja 외 arxiv

Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended externalities: utility vectors may weaken safety behavior, while refu…