paper-with-me

Papers

Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors

2025-12-13 · Veronica Mangiaterra, Hamad Al-Azary, Chiara Barattieri di San Pietro, Paolo Canal, Valentina Bambini arxiv

As Large Language Models (LLMs) are increasingly being used in scientific research, the issue of their trustworthiness becomes crucial. In psycholinguistics, LLMs have been recently employed in automatically augmenting human-rated datasets, with promising results obtained by generating ratings for single words. Yet, performance for ratings of complex items, i.e., metaphors, is still unexplored. Here, we present the first assessment of the validity and reliability of ratings of metaphors on familiarity, comprehensibility, and imageability, generated by three GPT models for a total of 687 items gathered from the Italian Figurative Archive and three English studies. We performed a thorough validation in terms of both alignment with human data and ability to predict behavioral and electrophysiological responses. We found that machine-generated ratings positively correlated with human-generated ones. Familiarity ratings reached moderate-to-strong correlations for both English and Italian metaphors, although correlations weakened for metaphors with high sensorimotor load. Imageability showed moderate correlations in English and moderate-to-strong in Italian. Comprehensibility for English metaphors exhibited the strongest correlations. Overall, larger models outperformed smaller ones and greater human-model misalignment emerged with familiarity and imageability. Machine-generated ratings significantly predicted response times and the EEG amplitude, with a strength comparable to human ratings. Moreover, GPT ratings obtained across independent sessions were highly stable. We conclude that GPT, especially larger models, can validly and reliably replace - or augment - human subjects in rating metaphor properties. Yet, LLMs align worse with humans when dealing with conventionality and multimodal aspects of metaphorical meaning, calling for careful consideration of the nature of stimuli.

📄 PDF Abstract BibTeX arXiv:2512.12444

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can ChatGPT and Bard Generate Aligned Assessment Items? A Reliability Analysis against Human Performance

2023-04-09 · Abdolvahab Khademi

ChatGPT and Bard are AI chatbots based on Large Language Models (LLM) that are slated to promise different applications in diverse areas. In education, these AI technologies have been tested for applications in assessmen…

Automated Essay Scoring

Automated Scoring of Graphical Open-Ended Responses Using Artificial Neural Networks

2022-01-05 · Matthias von Davier, Lillian Tyack, Lale Khorramdel

Automated scoring of free drawings or images as responses has yet to be utilized in large-scale assessments of student achievement. In this study, we propose artificial neural networks to classify these types of graphica…

Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory

2025-07-26 · Dan Song, Won-Chan Lee, Hong Jiao arxiv

This study investigates the estimation of reliability for large language models (LLMs) in scoring writing tasks from the AP Chinese Language and Culture Exam. Using generalizability theory, the research evaluates and com…

Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters

2025-10-29 · Xingjian Zhang, Tianhong Gao, Suliang Jin, Tianhao Wang 외 arxiv

Large language models (LLMs) are increasingly used as raters for evaluation tasks. However, their reliability is often limited for subjective tasks, when human judgments involve subtle reasoning beyond annotation labels.…

ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation

2024-05-08 · Ana Brassard, Benjamin Heinzerling, Keito Kudo, Keisuke Sakaguchi 외

Evaluating the quality of free-text explanations is a multifaceted, subjective, and labor-intensive task. Large language models (LLMs) present an appealing alternative due to their potential for consistency, scalability,…