paper-with-me

Papers

Tracing Persona Vectors Through LLM Pretraining

2026-05-13 · Viktor Moskvoretskii, Dominik Glandorf, Jorge Medina Moreira, Tanja Käser, Robert West arxiv

How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determines what we can detect, audit, or intervene on. Recent work has shown that traits such as evil or sycophancy correspond to linear directions in the internal activations, the so-called persona vectors. Although these vectors are now routinely utilized to inspect and steer model behavior in safety-relevant settings, how these representations are formed during training remains unknown. To address this gap, we trace persona vectors across the pretraining of OLMo-3-7B, finding that persona vectors form remarkably early -- within 0.22% of OLMo-3 pretraining -- and remain effective for steering the fully post-trained instruct models. Although core representations are formed early on, persona vectors continue to refine geometrically and semantically throughout pretraining. We further compare alternative elicitation strategies and find that all yield effective directions, with each strategy surfacing qualitatively distinct facets of the underlying persona. Replicating our analysis on Apertus-8B reveals that our findings transfer qualitatively beyond OLMo-3. Our results establish persona representations as stable features of early pretraining and open a path to studying how training forms, refines, and shapes them.

📄 PDF Abstract BibTeX arXiv:2605.13329

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HyperTrace: Hypothesis-Based Preference Tracing for Online LLM Personalization

2026-09-09 · Jianzhi Shen, Keyu Mao, Minghao Shao, Chuanyang Jin 외 arxiv

Personalized language models aim to adapt responses to individual users, whose preferences are often latent and revealed gradually through interaction. Existing training-free methods rely on stored histories or retrieved…

Tracing Multilingual Factual Knowledge Acquisition in Pretraining

2025-05-20 · Yihong Liu, Mingyang Wang, Amir Hossein Kargaran, Felicia Körner 외

Large Language Models (LLMs) are capable of recalling multilingual factual knowledge present in their pretraining data. However, most studies evaluate only the final model, leaving the development of factual recall and c…

Personalized Forgetting Mechanism with Concept-Driven Knowledge Tracing

2024-04-18 · Shanshan Wang, Ying Hu, Xun Yang, Zhongzhou Zhang 외

Knowledge Tracing (KT) aims to trace changes in students' knowledge states throughout their entire learning process by analyzing their historical learning data and predicting their future learning performance. Existing f…

Knowledge Tracing

Persona Vectors: Monitoring and Controlling Character Traits in Language Models

2025-07-29 · Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans 외 arxiv

Large language models interact with users through a simulated 'Assistant' persona. While the Assistant is typically trained to be helpful, harmless, and honest, it sometimes deviates from these ideals. In this paper, we …

Personal Devices for Contact Tracing: Smartphones and Wearables to Fight Covid-19

2021-08-02 · Pai Chet Ng, Petros Spachos, Stefano Gregori, Konstantinos Plataniotis

Digital contact tracing has emerged as a viable tool supplementing manual contact tracing. To date, more than 100 contact tracing applications have been published to slow down the spread of highly contagious Covid-19. De…