paper-with-me

Papers

Prompt Engineering for Scale Development in Generative Psychometrics

2026-03-16 · Lara Lee Russell-Lasalandra, Hudson Golino arxiv

This Monte Carlo simulation examines how prompt engineering strategies shape the quality of large language model (LLM)--generated personality assessment items within the AI-GENIE framework for generative psychometrics. Item pools targeting the Big Five traits were generated using multiple prompting designs (zero-shot, few-shot, persona-based, and adaptive), model temperatures, and LLMs, then evaluated and reduced using network psychometric methods. Across all conditions, AI-GENIE reliably improved structural validity following reduction, with the magnitude of its incremental contribution inversely related to the quality of the incoming item pool. Prompt design exerted a substantial influence on both pre- and post-reduction item quality. Adaptive prompting consistently outperformed non-adaptive strategies by sharply reducing semantic redundancy, elevating pre-reduction structural validity, and preserving substantially larger item pool, particularly when paired with newer, higher-capacity models. These gains were robust across temperature settings for most models, indicating that adaptive prompting mitigates common trade-offs between creativity and psychometric coherence. An exception was observed for the GPT-4o model at high temperatures, suggesting model-specific sensitivity to adaptive constraints at elevated stochasticity. Overall, the findings demonstrate that adaptive prompting is the strongest approach in this context, and that its benefits scale with model capability, motivating continued investigation of model--prompt interactions in generative psychometric pipelines.

📄 PDF Abstract BibTeX arXiv:2603.15909

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

Prompt Engineering for Responsible Generative AI Use in African Education: A Report from a Three-Day Training Series

2026-01-04 · Benjamin Quarshie, Vanessa Willemse, Macharious Nabang, Bismark Nyaaba Akanzire 외 arxiv

Generative artificial intelligence (GenAI) tools are increasingly adopted in education, yet many educators lack structured guidance on responsible and context sensitive prompt engineering, particularly in African and oth…

Prompt Engineering

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs

2026-06-02 · Nils Schwager, Christoph Hau, Simon Münker, Achim Rettinger arxiv

When prompting SLMs for psychometric assessments, researchers assume the outputs reflect semantic reasoning. We evaluate this premise across 13 open-weights models (0.6B to 14B parameters) using a prompt variation framew…

Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models

2024-09-18 · Haoran Ye, Yuhang Xie, Yuanyi Ren, Hanjun Fang 외

Human values and their measurement are long-standing interdisciplinary inquiry. Recent advances in AI have sparked renewed interest in this area, with large language models (LLMs) emerging as both tools and subjects of v…

Position: AI Evaluation Should Learn from How We Test Humans

2023-06-18 · Yan Zhuang, Qi Liu, Zachary A. Pardos, Patrick C. Kyllonen 외

As AI systems continue to evolve, their rigorous evaluation becomes crucial for their development and deployment. Researchers have constructed various large-scale benchmarks to determine their capabilities, typically aga…

Mathematical ReasoningPosition

Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design

2026-09-09 · Vinicius Kaster Marini, Petter Krus arxiv

The development of generative artificial intelligence resources enables opportunities of speeding up systems and engineering design work. This contribution introduces a framework of formal operations for assembling conte…