paper-with-me

Papers

Persona Vectors in Games: Measuring and Steering Strategies via Activation Vectors

2026-03-22 · Johnathan Sun, Andrew Zhang arxiv

Large language models (LLMs) are increasingly deployed as autonomous decision-makers in strategic settings, yet we have limited tools for understanding their high-level behavioral traits. We use activation steering methods in game-theoretic settings, constructing persona vectors for altruism, forgiveness, and expectations of others by contrastive activation addition. Evaluating on canonical games, we find that activation steering systematically shifts both quantitative strategic choices and natural-language justifications. However, we also observe that rhetoric and strategy can diverge under steering. In addition, vectors for self-behavior and expectations of others are partially distinct. Our results suggest that persona vectors offer a promising mechanistic handle on high-level traits in strategic environments.

📄 PDF Abstract BibTeX arXiv:2603.21398

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization

2024-05-28 · Yuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin 외

Researchers have been studying approaches to steer the behavior of Large Language Models (LLMs) and build personalized LLMs tailored for various applications. While fine-tuning seems to be a direct solution, it requires …

Hallucination

Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models

2026-01-23 · Pranav Bhandari, Usman Naseem, Mehwish Nasim arxiv

Personality steering in large language models (LLMs) commonly relies on injecting trait-specific steering vectors, implicitly assuming that personality traits can be controlled independently. In this work, we examine whe…

On the Limits of Steering Vectors for Preference-Aligned Generation

2026-07-02 · Melanie Subbiah, Zara Hall, Kathleen McKeown arxiv

Steering vectors have emerged as a promising approach to controlled text generation, offering interpretable, training-free mechanisms for shaping model outputs. However, their practical generality remains poorly understo…

Text Generation

Improving Steering Vectors by Targeting Sparse Autoencoder Features

2024-11-04 · Sviatoslav Chalnev, Matthew Siu, Arthur Conmy

To control the behavior of language models, steering methods attempt to ensure that outputs of the model satisfy specific pre-defined properties. Adding steering vectors to the model is a promising method of model contro…

Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy

2026-05-20 · Ishaan Kelkar, Nebras Alam, Vikram Kakaria, Madhur Panwar 외 arxiv

We study the effect of different persona on \textbf{sycophancy}: model's agreement with users even when the user is incorrect. The standard mitigation, Contrastive Activation Addition (CAA), derives a steering direction …