paper-with-me

홈 › Papers

Evaluating Language Model Character Traits

2024-10-05 · Francis Rhys Ward, Zejia Yang, Alex Jackson, Randy Brown, Chandler Smith, Grace Colverd, Louis Thomson, Raymond Douglas, Patrik Bartak, Andrew Rowan

Language models (LMs) can exhibit human-like behaviour, but it is unclear how to describe this behaviour without undue anthropomorphism. We formalise a behaviourist view of LM character traits: qualities such as truthfulness, sycophancy, or coherent beliefs and intentions, which may manifest as consistent patterns of behaviour. Our theory is grounded in empirical demonstrations of LMs exhibiting different character traits, such as accurate and logically coherent beliefs, and helpful and harmless intentions. We find that the consistency with which LMs exhibit certain character traits varies with model size, fine-tuning, and prompting. In addition to characterising LM character traits, we evaluate how these traits develop over the course of an interaction. We find that traits such as truthfulness and harmfulness can be stationary, i.e., consistent over an interaction, in certain contexts, but may be reflective in different contexts, meaning they mirror the LM's behavior in the preceding interaction. Our formalism enables us to describe LM behaviour precisely in intuitive language, without undue anthropomorphism.

📄 PDF Abstract BibTeX arXiv:2410.04272

Code (1)

graceebc9/agent_intentions 공식 구현

Tasks

Language ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires

2025-02-07 · Pranav Bhandari, Usman Naseem, Amitava Datta, Nicolas Fay 외

Psychological assessment tools have long helped humans understand behavioural patterns. While Large Language Models (LLMs) can generate content comparable to that of humans, we explore whether they exhibit personality tr…

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead

2025-07-30 · Tom Sühr, Florian E. Dorner, Olawale Salaudeen, Augustin Kelava 외 arxiv

Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive and psychological traits, such as intelligence and personality. While these res…

Orca: Enhancing Role-Playing Abilities of Large Language Models by Integrating Personality Traits

2024-11-15 · Yuxuan Huang

Large language models has catalyzed the development of personalized dialogue systems, numerous role-playing conversational agents have emerged. While previous research predominantly focused on enhancing the model's capab…

OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

2024-02-08 · Hainiu Xu, Runcong Zhao, Lixing Zhu, Jinhua Du 외

Neural Theory-of-Mind (N-ToM), machine's ability to understand and keep track of the mental states of others, is pivotal in developing socially intelligent agents. However, prevalent N-ToM benchmarks have several shortco…

Diversity

Can Language Models Recognize Convincing Arguments?

2024-03-31 · Paula Rescala, Manoel Horta Ribeiro, Tiancheng Hu, Robert West

The capabilities of large language models (LLMs) have raised concerns about their potential to create and propagate convincing narratives. Here, we study their performance in detecting convincing arguments to gain insigh…

Misinformation