paper-with-me

Papers

Contextual Linear Activation Steering of Language Models

2026-04-27 · Brandon Hsu, Daniel Beaglehole, Adityanarayanan Radhakrishnan, Mikhail Belkin arxiv

Linear activation steering is a powerful approach for eliciting the capabilities of large language models and specializing their behavior using limited labeled data. While effective, existing methods often apply a fixed steering strength to all tokens, resulting in inconsistent steering quality across diverse input prompts. In this work, we introduce Contextual Linear Activation Steering (CLAS), a method that dynamically adapts linear activation steering to context-dependent steering strengths. Across eleven steering benchmarks and four model families, it consistently outperforms standard linear activation steering and matches or exceeds the performance of ReFT and LoRA in settings with limited labeled data. We therefore propose CLAS as a scalable, interpretable, and accurate method for specializing and steering large language models.

📄 PDF Abstract BibTeX arXiv:2604.24693

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Curveball Steering: The Right Direction To Steer Isn't Always Linear

2026-03-10 · Shivam Raval, Hae Jin Song, Linlin Wu, Abir Harrasse 외 arxiv

Activation steering is a widely used approach for controlling large language model (LLM) behavior by intervening on internal representations. Existing methods largely rely on the Linear Representation Hypothesis, assumin…

Beyond Linear Activation Steering: Invertible Latent Transformations for Controlling LLM Behavior

2026-06-07 · Tuc Nguyen, Thai Le arxiv

Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors. Most existing methods compute a fi…

Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

2026-03-31 · Haoran Wang, Li Xiong, Kai Shu arxiv

Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they frequently violate contextual privacy by disclosing private information in situations where humans would exercise discretion. This …

Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations

2026-02-19 · Joschka Braun arxiv

Steering vectors are a lightweight method for controlling language model behavior by adding a learned bias to the activations at inference time. Although effective on average, steering effect sizes vary across samples an…

Beyond Linear Steering: Unified Multi-Attribute Control for Language Models

2025-05-30 · Narmeen Oozeer, Luke Marks, Fazl Barez, Amirali Abdullah

Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of linear steering methods, which assume ad…

Attribute