paper-with-me

Papers

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models

2026-04-24 · Juan Manuel Contreras arxiv

Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models behave. Is this gap an artifact of forcing human trait categories onto LLMs, or something deeper about LLM self-report? To find out, we built the first psychometric instrument whose dimensions are derived from LLM behavior rather than human psychology. Administering 300 items (240 Likert + 60 scenario) to 25 LLMs across 17 model families, 30 times each, exploratory factor analysis revealed five reliable, replicable factors: Responsiveness, Deference, Boldness, Guardedness, and Verbosity (all Tucker $φ\geq .957$, all $α\geq .930$). We collected 2,500 open-ended samples and had them rated by 151 humans and a three-judge LLM ensemble. Humans and judges agreed ($\bar{r} = .51$), but self-report predicted neither the ratings nor objective text measures computed from them: the gap persists even for constructs native to LLMs, where a human-mismatch explanation no longer applies. The exception is Verbosity, whose self-report reaches 74% of the criterion-reliability ceiling against human ratings, but does not track raw output length. On Responsiveness, self-report tracked LLM judges ($r = .53$) but not humans ($r = .04$), even though humans and judges otherwise agreed ($r = .59$). This pattern formally rejects any single latent construct driving all three measurements ($p = .007$). Self-report items and LLM judges share a source of variance that human observers do not, and controlling for measurable surface features (length, formatting, enthusiasm markers) does not remove it. This confound is invisible to the within-ensemble reliability checks used to validate LLM judges, and it poses a concrete risk for the LLM-as-judge pipelines now central to model evaluation. We release the instrument as a diagnostic probe for alignment-shaped self-description.

📄 PDF Abstract BibTeX arXiv:2606.09843

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A validity-guided workflow for robust large language model research in psychology

2025-07-06 · Zhicheng Lin arxiv

Large language models (LLMs) are rapidly being integrated into psychological research as research tools, evaluation targets, human simulators, and cognitive models. However, recent evidence reveals severe measurement unr…

Causal Inference

Perla: A Conversational Agent for Depression Screening in Digital Ecosystems. Design, Implementation and Validation

2020-08-28 · Raúl Arrabales

Most depression assessment tools are based on self-report questionnaires, such as the Patient Health Questionnaire (PHQ-9). These psychometric instruments can be easily adapted to an online setting by means of electronic…

Specificity

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

2025-10-25 · Alexandra Yost, Shreyans Jain, Shivam Raval, Grant Corser 외 arxiv

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We propose a framework to measure consistent be…

Large Language Models as Simulative Agents for Neurodivergent Adult Psychometric Profiles

2026-01-16 · Francesco Chiappone, Davide Marocco, Nicola Milano arxiv

Adult neurodivergence, including Attention-Deficit/Hyperactivity Disorder (ADHD), high-functioning Autism Spectrum Disorder (ASD), and Cognitive Disengagement Syndrome (CDS), is marked by substantial symptom overlap that…

Items from Psychometric Tests as Training Data for Personality Profiling Models of Twitter Users

2022-02-21 · WASSA (ACL) 2022 5 · Anne Kreuter, Kai Sassenberg, Roman Klinger

Machine-learned models for author profiling in social media often rely on data acquired via self-reporting-based psychometric tests (questionnaires) filled out by social media users. This is an expensive but accurate dat…

Author ProfilingData Augmentation