paper-with-me

홈 › Papers

AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans

2025-09-20 · Wei Xie, Shuoyoucheng Ma, Zhenhua Wang, Enze Wang, Kai Chen, Xiaobing Sun, Baosheng Wang arxiv

Large Language Models (LLMs) with hundreds of billions of parameters have exhibited human-like intelligence by learning from vast amounts of internet-scale data. However, the uninterpretability of large-scale neural networks raises concerns about the reliability of LLM. Studies have attempted to assess the psychometric properties of LLMs by borrowing concepts from human psychology to enhance their interpretability, but they fail to account for the fundamental differences between LLMs and humans. This results in high rejection rates when human scales are reused directly. Furthermore, these scales do not support the measurement of LLM psychological property variations in different languages. This paper introduces AIPsychoBench, a specialized benchmark tailored to assess the psychological properties of LLM. It uses a lightweight role-playing prompt to bypass LLM alignment, improving the average effective response rate from 70.12% to 90.40%. Meanwhile, the average biases are only 3.3% (positive) and 2.1% (negative), which are significantly lower than the biases of 9.8% and 6.9%, respectively, caused by traditional jailbreak prompts. Furthermore, among the total of 112 psychometric subcategories, the score deviations for seven languages compared to English ranged from 5% to 20.2% in 43 subcategories, providing the first comprehensive evidence of the linguistic impact on the psychometrics of LLM.

📄 PDF Abstract BibTeX arXiv:2509.16530

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Machine individuality: Separating genuine idiosyncrasy from response bias in large language models

2026-04-18 · Valentin Kriegmair, Dirk U. Wulff arxiv

As large language models (LLMs) are increasingly integrated into daily life, in roles ranging from high-stakes decision support to companionship, understanding their behavioral dispositions becomes critical. A growing li…

Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias

2026-01-08 · Adib Sakhawat, Tahsin Islam, Takia Farhin, Syed Rifat Raiyan 외 arxiv

As large language models (LLMs) are increasingly deployed, understanding how they express political positioning is important for evaluating alignment and downstream effects. We audit 26 contemporary LLMs using three poli…

Quantifying AI Psychology: A Psychometrics Benchmark for Large Language Models

2024-06-25 · Yuan Li, Yue Huang, Hongyi Wang, Xiangliang Zhang 외

Large Language Models (LLMs) have demonstrated exceptional task-solving capabilities, increasingly adopting roles akin to human-like assistants. The broader integration of LLMs into society has sparked interest in whethe…

Cognitive phantoms in LLMs through the lens of latent variables

2024-09-06 · Sanne Peereboom, Inga Schwabe, Bennett Kleinberg

Large language models (LLMs) increasingly reach real-world applications, necessitating a better understanding of their behaviour. Their size and complexity complicate traditional assessment methods, causing the emergence…

Stories of Your Life as Others: A Round-Trip Evaluation of LLM-Generated Life Stories Conditioned on Rich Psychometric Profiles

2026-04-07 · Ben Wigler, Maria Tsfasman, Tiffany Matej Hrkalovic arxiv

Personality traits are richly encoded in natural language, and large language models (LLMs) trained on human text can simulate personality when conditioned on persona descriptions. However, existing evaluations rely pred…