paper-with-me

홈 › Papers

A validity-guided workflow for robust large language model research in psychology

2025-07-06 · Zhicheng Lin arxiv

Large language models (LLMs) are rapidly being integrated into psychological research as research tools, evaluation targets, human simulators, and cognitive models. However, recent evidence reveals severe measurement unreliability: Personality assessments collapse under factor analysis, moral preferences reverse with punctuation changes, and theory-of-mind accuracy varies widely with trivial rephrasing. These "measurement phantoms"--statistical artifacts masquerading as psychological phenomena--threaten the validity of a growing body of research. Guided by the dual-validity framework that integrates psychometrics with causal inference, we present a six-stage workflow that scales validity requirements to research ambition--using LLMs to code text requires basic reliability and accuracy, while claims about psychological properties demand comprehensive construct validation. Researchers must (1) explicitly define their research goal and corresponding validity requirements, (2) develop and validate computational instruments through psychometric testing, (3) design experiments that control for computational confounds, (4) execute protocols with transparency, (5) analyze data using methods appropriate for non-independent observations, and (6) report findings within demonstrated boundaries and use results to refine theory. We illustrate the workflow through an example of model evaluation--"LLM selfhood"--showing how systematic validation can distinguish genuine computational phenomena from measurement artifacts. By establishing validated computational instruments and transparent practices, this workflow provides a path toward building a robust empirical foundation for AI psychology research.

📄 PDF Abstract BibTeX arXiv:2507.04491

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Inference

Similar Papers 제목 키워드 기반

Agentic AI for Remote Sensing: Technical Challenges and Research Directions

2026-04-27 · Muhammad Akhtar Munir, Muhammad Umer Sheikh, Akashah Shabbir, Muhammad Haris Khan 외 arxiv

Earth Observation (EO) is moving beyond static prediction toward multi-step analytical workflows that require coordinated reasoning over data, tools, and geospatial state. While foundation models and vision-language mode…

Representation Learning

GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation

2025-09-07 · Qianheng Zhang, Song Gao, Chen Wei, Yibo Zhao 외 arxiv

Recent advances in large language models (LLMs) have fueled growing interest in automating geospatial analysis and GIS workflows, yet their actual capabilities remain uncertain. In this work, we call for rigorous evaluat…

Semantic SimilaritySpatial ReasoningCode Generation

ATLAS: A Layered Constraint-Guided Framework for Structured Artifact Generation in LLM-Assisted MDE

2025-10-29 · Tong Ma, Hui Lai, Hui Wang, Zhenhu Tian 외 arxiv

ATLAS is a constraint-guided generation framework for structured engineering artifacts whose outputs must satisfy explicit schemas, domain rules, and audit requirements. Rather than treating a large language model as a s…

Agentic Workflows for Economic Research: Design and Implementation

2025-04-13 · Herbert Dawid, Philipp Harting, Hankui Wang, Zhongli Wang 외

This paper introduces a methodology based on agentic workflows for economic research that leverages Large Language Models (LLMs) and multimodal AI to enhance research efficiency and reproducibility. Our approach features…

ComfyUI-R1: Exploring Reasoning Models for Workflow Generation

2025-06-11 · Zhenran Xu, Yiyu Wang, Xue Yang, Longyue Wang 외

AI-generated content has evolved from monolithic models to modular workflows, particularly on platforms like ComfyUI, enabling customization in creative pipelines. However, crafting effective workflows requires great exp…

4k