Patterns and Mechanisms of Contrastive Activation Engineering
Controlling the behavior of Large Language Models (LLMs) remains a significant challenge due to their inherent complexity and opacity. While techniques like fine-tuning can modify model behavior, they typically require extensive computational resources. Recent work has introduced a class of contrastive activation engineering (CAE) techniques as promising approaches for steering LLM outputs through targeted modifications to their internal representations. Applied at inference-time with zero cost, CAE has the potential to introduce a new paradigm of flexible, task-specific LLM behavior tuning. We analyze the performance of CAE in in-distribution, out-of-distribution settings, evaluate drawbacks, and begin to develop comprehensive guidelines for its effective deployment. We find that 1. CAE is only reliably effective when applied to in-distribution contexts. 2. Increasing the number of samples used to generate steering vectors has diminishing returns at around 80 samples. 3. Steering vectors are susceptible to adversarial inputs that reverses the behavior that is steered for. 4. Steering vectors harm the overall model perplexity. 5. Larger models are more resistant to steering-induced degradation.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering
The field of large language models (LLMs) has grown rapidly in recent years, driven by the desire for better efficiency, interpretability, and safe use. Building on the novel approach of "activation engineering," this st…
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding.However, hallucination, where the model generates plausible yet incorrect outputs, persists as a significant and under-add…
HallucinationRevisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
The recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these th…
Prompt EngineeringUnderstanding the Role of Nonlinearity in Training Dynamics of Contrastive Learning
While the empirical success of self-supervised learning (SSL) heavily relies on the usage of deep nonlinear models, existing theoretical works on SSL understanding still focus on linear ones. In this paper, we study the …
Contrastive LearningSelf-Supervised LearningOncolytic mechanisms and immunotherapeutic potential of Newcastle disease virus in cancer therapy
Newcastle Disease Virus (NDV), classified as Avian orthoavulavirus 1 (avian paramyxovirus type 1), is a promising oncolytic agent that selectively targets and destroys cancer cells while sparing normal tissues. Its oncos…
Morphology classification