paper-with-me

홈 › Papers

Can LLMs Infer Personality from Real World Conversations?

2025-07-18 · Jianfeng Zhu, Ruoming Jin, Karin G. Coifman arxiv

Large Language Models (LLMs) such as OpenAI's GPT-4 and Meta's LLaMA offer a promising approach for scalable personality assessment from open-ended language. However, inferring personality traits remains challenging, and earlier work often relied on synthetic data or social media text lacking psychometric validity. We introduce a real-world benchmark of 555 semi-structured interviews with BFI-10 self-report scores for evaluating LLM-based personality inference. Three state-of-the-art LLMs (GPT-4.1 Mini, Meta-LLaMA, and DeepSeek) were tested using zero-shot prompting for BFI-10 item prediction and both zero-shot and chain-of-thought prompting for Big Five trait inference. All models showed high test-retest reliability, but construct validity was limited: correlations with ground-truth scores were weak (max Pearson's $r = 0.27$), interrater agreement was low (Cohen's $κ< 0.10$), and predictions were biased toward moderate or high trait levels. Chain-of-thought prompting and longer input context modestly improved distributional alignment, but not trait-level accuracy. These results underscore limitations in current LLM-based personality inference and highlight the need for evidence-based development for psychological applications.

📄 PDF Abstract BibTeX arXiv:2507.14355

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind

2025-01-15 · Kazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno 외

Existing Theory of Mind (ToM) benchmarks diverge from real-world scenarios in three aspects: 1) they assess a limited range of mental states such as beliefs, 2) false beliefs are not comprehensively explored, and 3) the …

BenchmarkingMultiple-choice

Investigating Large Language Models in Inferring Personality Traits from User Conversations

2025-01-13 · Jianfeng Zhu, Ruoming Jin, Karin G. Coifman

Large Language Models (LLMs) are demonstrating remarkable human like capabilities across diverse domains, including psychological assessment. This study evaluates whether LLMs, specifically GPT-4o and GPT-4o mini, can in…

Orca: Enhancing Role-Playing Abilities of Large Language Models by Integrating Personality Traits

2024-11-15 · Yuxuan Huang

Large language models has catalyzed the development of personalized dialogue systems, numerous role-playing conversational agents have emerged. While previous research predominantly focused on enhancing the model's capab…

Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution

2026-02-07 · Deuksin Kwon, Kaleen Shrestha, Bin Han, Spencer Lin 외 arxiv

Large language models (LLMs) are increasingly used to simulate human behavior in social settings such as legal mediation, negotiation, and dispute resolution. However, it remains unclear whether these simulations reprodu…

Modeling Dyadic Conversations for Personality Inference

2020-09-26 · Qiang Liu

Nowadays, automatical personality inference is drawing extensive attention from both academia and industry. Conventional methods are mainly based on user generated contents, e.g., profiles, likes, and texts of an individ…