paper-with-me

Papers

What Does Activation Steering Control? Attribution Across Answer Encodings and Output-Sensitive Subspaces

2026-08-24 · Zhiwei Gao, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki arxiv

Activation steering is often evaluated under the answer encoding used to construct the direction. A reported gain may reflect the intended judgment or compatibility with answer identifiers seen during construction. We introduce Cross-Encoding Steering Evaluation, which freezes an intervention while re-encoding answers to the same held-out items. On NormBank, after A/B/C identifiers are reassigned, contrastive activation addition (CAA) induces larger target-versus-source score changes for the extraction indices than for the semantic labels under the new mapping. We call this extraction-index following. Varying identifier vocabulary (A/B/C, X/Y/Z, or 1/2/3) and row order shows that the effect tracks extraction index rather than row position. After matching direction norms across layers, extraction-index following emerges mainly at later depths. A low-rank output-sensitive component containing 15.4% of the direction's squared norm retains 96.3% of this effect. An Inference-Time Intervention (ITI)-style method also favors extraction-index over semantic-label following on NormBank in three models. In aggregate, MNLI favors extraction-index following, whereas Social Chemistry 101 (SC101) favors semantic-label following. Multiple-choice and open-ended evaluations can yield different behavioral conclusions. Thus, a steering gain under one answer encoding does not by itself identify what the intervention controls.

📄 PDF Abstract BibTeX arXiv:2608.22985

Code (0)

등록된 구현이 없습니다.

Tasks

Steering Control

Similar Papers 제목 키워드 기반

Decomposing Theory of Mind: How Emotional Processing Mediates ToM Abilities in LLMs

2025-11-19 · Ivan Chulo, Ananya Joshi arxiv

Recent work shows activation steering substantially improves language models' Theory of Mind (ToM) (Bortoletto et al. 2024), yet the mechanisms of what changes occur internally that leads to different outputs remains unc…

Where Steering Signals Come From: Activation Source Selection in Activation Steering

2026-07-28 · Jiaran Ye, Lingxu Ran, Zijun Yao, Chenpeng Wang 외 arxiv

Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail. We study this sourc…

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs

2025-07-24 · Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal arxiv

Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.…

What Can We Actually Steer? A Multi-Behavior Study of Activation Control

2025-11-23 · Tetiana Bas, Krystian Novak arxiv

Large language models (LLMs) require precise behavior control for safe and effective deployment across diverse applications. Activation steering offers a promising approach for LLMs' behavioral control. We focus on the q…

From Attribution to Action: A Human-Centered Application of Activation Steering

2026-04-13 · Tobias Labarta, Maximilian Dreyer, Katharina Weitz, Wojciech Samek 외 arxiv

Explainable AI (XAI) methods reveal which features influence model predictions, yet provide limited means for practitioners to act on these explanations. Activation steering of components identified via XAI offers a path…