BigFive: A Dataset of Coarse- and Fine-Grained Personality Characteristics
Obtaining the personalities of users conveyed by their published short texts has a wide and important range of applications, from detecting abnormal behavior of online users to accurately customization recommendation. Advancement in this area can be improved using large-scale datasets with coarse- and fine-grained typologies, adaptable to multiple downstream tasks. Therefore, this paper introduces $BigFive$, a large, high quality dataset manually annotated by experts. $BigFive$ contains 13,478 Chinese phrases that belong to five categories (coarse-grained) and 30 categories (fine-grained). The reliability of five categories grouped by personality level and 30 categories grouped by dimension level is demonstrated via a detailed data analysis. In addition, a strong baseline is build based on fine-tuning a BERT model. Our BERT-based model achieves an average F1-score of .33 (std=.24) in terms of 30 categories and an average F1-score of .66 (std=.05) in terms of five categories. The experimental results suggest that there is much room for improvement.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
PsyAttention: Psychological Attention Model for Personality Detection
Work on personality detection has tended to incorporate psychological features from different personality models, such as BigFive and MBTI. There are more than 900 psychological features, each of which is helpful for per…
modelLITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space
Personalized text generation for authors and literary writing is essential for applications such as adaptive writing assistants, creative support tools, and computational literary analysis. However, existing approaches t…
Text GenerationPredicting the Big Five Personality Traits in Chinese Counselling Dialogues Using Large Language Models
Accurate assessment of personality traits is crucial for effective psycho-counseling, yet traditional methods like self-report questionnaires are time-consuming and biased. This study exams whether Large Language Models …
Orca: Enhancing Role-Playing Abilities of Large Language Models by Integrating Personality Traits
Large language models has catalyzed the development of personalized dialogue systems, numerous role-playing conversational agents have emerged. While previous research predominantly focused on enhancing the model's capab…
Facet-Level Persona Control by Trait-Activated Routing with Contrastive SAE for Role-Playing LLMs
Personality control in Role-Playing Agents (RPAs) is commonly achieved via training-free methods that inject persona descriptions and memory through prompts or retrieval-augmented generation, or via supervised fine-tunin…