paper-with-me

Papers

Identifying and Manipulating the Personality Traits of Language Models

2022-12-20 · Graham Caron, Shashank Srivastava

Psychology research has long explored aspects of human personality such as extroversion, agreeableness and emotional stability. Categorizations like the Big Five' personality traits are commonly used to assess and diagnose personality types. In this work, we explore the question of whether the perceived personality in language models is exhibited consistently in their language generation. For example, is a language model such as GPT2 likely to respond in a consistent way if asked to go out to a party? We also investigate whether such personality traits can be controlled. We show that when provided different types of contexts (such as personality descriptions, or answers to diagnostic questions about personality traits), language models such as BERT and GPT2 can consistently identify and reflect personality markers in those contexts. This behavior illustrates an ability to be manipulated in a highly predictable way, and frames them as tools for identifying personality traits and controlling personas in applications such as dialog systems. We also contribute a crowd-sourced data-set of personality descriptions of human subjects paired with their Big Five' personality assessment data, and a data-set of personality descriptions collated from Reddit.

📄 PDF Abstract BibTeX arXiv:2212.10276

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticLanguage ModelingLanguage ModellingText Generation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Residual Connection 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering

2024-12-10 · Rumi A. Allbert, James K. Wiles, Vlad Grankovsky

The field of large language models (LLMs) has grown rapidly in recent years, driven by the desire for better efficiency, interpretability, and safe use. Building on the novel approach of "activation engineering," this st…

Identifying Personality Traits Using Overlap Dynamics in Multiparty Dialogue

2019-09-02 · Mingzhi Yu, Emer Gilmartin, Diane Litman

Research on human spoken language has shown that speech plays an important role in identifying speaker personality traits. In this work, we propose an approach for identifying speaker personality traits using overlap dyn…

Neuron-based Personality Trait Induction in Large Language Models

2024-10-16 · Jia Deng, Tianyi Tang, Yanbin Yin, Wenhao Yang 외

Large language models (LLMs) have become increasingly proficient at simulating various personality traits, an important capability for supporting related applications (e.g., role-playing). To further improve this capacit…

Effects of personality steering on cooperative behavior in Large Language Model agents

2026-01-08 · Mizuki Sakai, Mizuki Yokoyama, Wakaba Tateishi, Genki Ichinose arxiv

Large language models (LLMs) are increasingly used as autonomous agents in strategic and social interactions. Although recent studies suggest that assigning personality traits to LLMs can influence their behavior, how pe…

When Does Personality Composition Matter for Multi-Agent LLM Teams?

2026-06-25 · Aryan Keluskar, Amrita Bhattacharjee, Huan Liu arxiv

Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-explored. Prior work shows that agents prompted with low agreeableness …