paper-with-me

홈 › Papers

The Geometry of Personality: Activation Steering with Jungian Cognitive Functions

2026-07-23 · Liu Zai, Yumeng Wang, Junchen Fu, Joemon M. Jose arxiv

Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled as a set of cognitive processes using the eight Jungian Cognitive Functions. To this end, we introduce a framework comprising a Jungian evaluation protocol and a dataset of over 2,100 role-playing character narrations. Activation steering vector extraction and evaluation experiments on Llama-3.1-8B demonstrate effective monotonic control over all eight cognitive functions through activation steering. Beyond controllability, our analysis reveals that: 1. personality information is concentrated in middle transformer layers; 2. steering vectors exhibit structured geometric relationships consistent with distinctions between rational and irrational functions; 3. effective multi-dimensional steering directions cannot be recovered as linear combinations of single-function directions. These findings provide new insights into the representation of personality in LLM activation space and establish a framework for studying interpretable, effective, and multi-dimensional personality control.

📄 PDF Abstract BibTeX arXiv:2607.20803

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 112
arxivsub/arXivSub_daily_arxiv ★ 3

Similar Papers 제목 키워드 기반

Bayesian algorithmic perfumery: A Hierarchical Relevance Vector Machine for the Estimation of Personalized Fragrance Preferences based on Three Sensory Layers and Jungian Personality Archetypes

2024-11-06 · Rolando Gonzales Martinez

This study explores a Bayesian algorithmic approach to personalized fragrance recommendation by integrating hierarchical Relevance Vector Machines (RVM) and Jungian personality archetypes. The paper proposes a structured…

Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

2026-06-27 · David Courtis, Ting Hu arxiv

Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. I…

Prompt Engineering

Linear Personality Probing and Steering in LLMs: A Big Five Study

2025-12-19 · Michel Frising, Daniel Balcells arxiv

Large language models (LLMs) exhibit distinct and consistent personalities that greatly impact trust and engagement. While this means that personality frameworks would be highly valuable tools to characterize and control…

Prompt Engineering

Controllable and explainable personality sliders for LLMs at inference time

2026-02-10 · Florian Hoppe, David Khachaturov, Robert Mullins, Mark Huasong Meng arxiv

Aligning Large Language Models (LLMs) with specific personas typically relies on expensive and monolithic Supervised Fine-Tuning (SFT) or RLHF. While effective, these methods require training distinct models for every ta…

Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering

2024-12-10 · Rumi A. Allbert, James K. Wiles, Vlad Grankovsky

The field of large language models (LLMs) has grown rapidly in recent years, driven by the desire for better efficiency, interpretability, and safe use. Building on the novel approach of "activation engineering," this st…