paper-with-me

홈 › Papers

Patterns and Mechanisms of Contrastive Activation Engineering

2025-05-06 · Yixiong Hao, Ayush Panda, Stepan Shabalin, Sheikh Abdur Raheem Ali

Controlling the behavior of Large Language Models (LLMs) remains a significant challenge due to their inherent complexity and opacity. While techniques like fine-tuning can modify model behavior, they typically require extensive computational resources. Recent work has introduced a class of contrastive activation engineering (CAE) techniques as promising approaches for steering LLM outputs through targeted modifications to their internal representations. Applied at inference-time with zero cost, CAE has the potential to introduce a new paradigm of flexible, task-specific LLM behavior tuning. We analyze the performance of CAE in in-distribution, out-of-distribution settings, evaluate drawbacks, and begin to develop comprehensive guidelines for its effective deployment. We find that 1. CAE is only reliably effective when applied to in-distribution contexts. 2. Increasing the number of samples used to generate steering vectors has diminishing returns at around 80 samples. 3. Steering vectors are susceptible to adversarial inputs that reverses the behavior that is steered for. 4. Steering vectors harm the overall model perplexity. 5. Larger models are more resistant to steering-induced degradation.

📄 PDF Abstract BibTeX arXiv:2505.03189

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering

2024-12-10 · Rumi A. Allbert, James K. Wiles, Vlad Grankovsky

The field of large language models (LLMs) has grown rapidly in recent years, driven by the desire for better efficiency, interpretability, and safe use. Building on the novel approach of "activation engineering," this st…

Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering

2025-05-19 · JianFeng Cai, Wengang Zhou, Zongmeng Zhang, Jiale Hong 외

Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding.However, hallucination, where the model generates plausible yet incorrect outputs, persists as a significant and under-add…

Hallucination

Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

2024-01-12 · Tianlong Li, Zhenghua Wang, Wenhao Liu, Muling Wu 외

The recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) when exposed to malicious inputs. While various defense strategies have been proposed to mitigate these th…

Prompt Engineering

Understanding the Role of Nonlinearity in Training Dynamics of Contrastive Learning

2022-06-02 · Yuandong Tian

While the empirical success of self-supervised learning (SSL) heavily relies on the usage of deep nonlinear models, existing theoretical works on SSL understanding still focus on linear ones. In this paper, we study the …

Contrastive LearningSelf-Supervised Learning

Oncolytic mechanisms and immunotherapeutic potential of Newcastle disease virus in cancer therapy

2025-05-09 · Umar Ahmad, Surializa Harun, Moussa Moise Diagne, Syahril Abdullah 외

Newcastle Disease Virus (NDV), classified as Avian orthoavulavirus 1 (avian paramyxovirus type 1), is a promising oncolytic agent that selectively targets and destroys cancer cells while sparing normal tissues. Its oncos…

Morphology classification