paper-with-me

홈 › Papers

R.U.Psycho? Robust Unified Psychometric Testing of Language Models

2025-03-13 · Julian Schelb, Orr Borin, David Garcia, Andreas Spitz

Generative language models are increasingly being subjected to psychometric questionnaires intended for human testing, in efforts to establish their traits, as benchmarks for alignment, or to simulate participants in social science experiments. While this growing body of work sheds light on the likeness of model responses to those of humans, concerns are warranted regarding the rigour and reproducibility with which these experiments may be conducted. Instabilities in model outputs, sensitivity to prompt design, parameter settings, and a large number of available model versions increase documentation requirements. Consequently, generalization of findings is often complex and reproducibility is far from guaranteed. In this paper, we present R.U.Psycho, a framework for designing and running robust and reproducible psychometric experiments on generative language models that requires limited coding expertise. We demonstrate the capability of our framework on a variety of psychometric questionnaires, which lend support to prior findings in the literature. R.U.Psycho is available as a Python package at https://github.com/julianschelb/rupsycho.

📄 PDF Abstract BibTeX arXiv:2503.10229

Code (1)

julianschelb/rupsycho 공식 구현

Similar Papers 제목 키워드 기반

Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality

2025-10-13 · Jana Jung, Marlene Lutz, Indira Sen, Markus Strohmaier arxiv

Psychometric tests are increasingly used to assess psychological constructs in large language models (LLMs). However, it remains unclear whether these tests -- originally developed for humans -- yield meaningful results …

Predicting Human Psychometric Properties Using Computational Language Models

2022-05-12 · Antonio Laverghetta Jr., Animesh Nighojkar, Jamshidbek Mirzakhalov, John Licato

Transformer-based language models (LMs) continue to achieve state-of-the-art performance on natural language processing (NLP) benchmarks, including tasks designed to mimic human-inspired "commonsense" competencies. To be…

Diagnostic

AI Psychometrics: Evaluating the Psychological Reasoning of Large Language Models with Psychometric Validities

2026-03-11 · Yibai Li, Xiaolin Lin, Zhenghui Sha, Zhiye Jin 외 arxiv

The immense number of parameters and deep neural networks make large language models (LLMs) rival the complexity of human brains, which also makes them opaque ``black box'' systems that are challenging to evaluate and in…

Constructing a Testbed for Psychometric Natural Language Processing

2020-07-25 · Ahmed Abbasi, David G. Dobolyi, Richard G. Netemeyer

Psychometric measures of ability, attitudes, perceptions, and beliefs are crucial for understanding user behaviors in various contexts including health, security, e-commerce, and finance. Traditionally, psychometric dime…

Survey

Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement

2025-05-13 · Haoran Ye, Jing Jin, Yuhang Xie, Xin Zhang 외

The rapid advancement of large language models (LLMs) has outpaced traditional evaluation methodologies. It presents novel challenges, such as measuring human-like psychological constructs, navigating beyond static and t…

BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model