paper-with-me

홈 › Papers

Automated Item Neutralization for Non-Cognitive Scales: A Large Language Model Approach to Reducing Social-Desirability Bias

2025-09-09 · Sirui Wu, Daijin Yang arxiv

This study evaluates item neutralization assisted by the large language model (LLM) to reduce social desirability bias in personality assessment. GPT-o3 was used to rewrite the International Personality Item Pool Big Five Measure (IPIP-BFM-50), and 203 participants completed either the original or neutralized form along with the Marlowe-Crowne Social Desirability Scale. The results showed preserved reliability and a five-factor structure, with gains in Conscientiousness and declines in Agreeableness and Openness. The correlations with social desirability decreased for several items, but inconsistently. Configural invariance held, though metric and scalar invariance failed. Findings support AI neutralization as a potential but imperfect bias-reduction method.

📄 PDF Abstract BibTeX arXiv:2509.19314

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings

2025-10-30 · Andrew M. Bean, Nabeel Seedat, Shengzhuang Chen, Jonathan Richard Schwarz arxiv

The prohibitive cost of evaluating large language models (LLMs) on comprehensive benchmarks necessitates the creation of small yet representative data subsets (i.e., tiny benchmarks) that enable efficient assessment whil…

LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?")

2026-06-03 · Alvin Wei Ming Tan, David Cardinal, Tania Lorido-Botran, Laura Bravo-Sanchez 외 arxiv

Given the inherently multimodal nature of human experience, vision-language models (VLMs) hold substantial promise for modeling human cognition as it grows and develops with experience. Realizing their potential requires…

Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report

2026-04-20 · Jon-Paul Cacioli arxiv

Clinical personality assessment screens response validity before interpreting substantive scales. LLM evaluation does not. We apply the validity scaling framework from the PAI and MMPI-3 to metacognitive probe data from …

Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?

2025-10-29 · Seonjeong Hwang, Hyounghun Kim, Gary Geunbae Lee arxiv

Estimating the cognitive complexity of reading comprehension (RC) items is crucial for assessing item difficulty before it is administered to learners. Unlike syntactic and semantic features, such as passage length or se…

Reading ComprehensionSemantic Similarity

Cognitively Diverse Multiple-Choice Question Generation: A Hybrid Multi-Agent Framework with Large Language Models

2026-02-03 · Yu Tian, Linh Huynh, Katerina Christhilf, Shubham Chakraborty 외 arxiv

Recent advances in large language models (LLMs) have made automated multiple-choice question (MCQ) generation increasingly feasible; however, reliably producing items that satisfy controlled cognitive demands remains a c…

Reading ComprehensionQuestion Generation