paper-with-me

Papers

How do Humans and Language Models Reason About Creativity? A Comparative Analysis

2025-02-05 · Antonio Laverghetta Jr., Tuhin Chakrabarty, Tom Hope, Jimmy Pronchick, Krupa Bhawsar, Roger E. Beaty

Creativity assessment in science and engineering is increasingly based on both human and AI judgment, but the cognitive processes and biases behind these evaluations remain poorly understood. We conducted two experiments examining how including example solutions with ratings impact creativity evaluation, using a finegrained annotation protocol where raters were tasked with explaining their originality scores and rating for the facets of remoteness (whether the response is "far" from everyday ideas), uncommonness (whether the response is rare), and cleverness. In Study 1, we analyzed creativity ratings from 72 experts with formal science or engineering training, comparing those who received example solutions with ratings (example) to those who did not (no example). Computational text analysis revealed that, compared to experts with examples, no-example experts used more comparative language (e.g., "better/worse") and emphasized solution uncommonness, suggesting they may have relied more on memory retrieval for comparisons. In Study 2, parallel analyses with state-of-the-art LLMs revealed that models prioritized uncommonness and remoteness of ideas when rating originality, suggesting an evaluative process rooted around the semantic similarity of ideas. In the example condition, while LLM accuracy in predicting the true originality scores improved, the correlations of remoteness, uncommonness, and cleverness with originality also increased substantially -- to upwards of $0.99$ -- suggesting a homogenization in the LLMs evaluation of the individual facets. These findings highlight important implications for how humans and AI reason about creativity and suggest diverging preferences for what different populations prioritize when rating.

📄 PDF Abstract BibTeX arXiv:2502.03253

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

A Comparative Approach to Assessing Linguistic Creativity of Large Language Models and Humans

2025-07-16 · Anca Dinu, Andra-Maria Florescu, Alina Resceanu arxiv

The following paper introduces a general linguistic creativity test for humans and Large Language Models (LLMs). The test consists of various tasks aimed at assessing their ability to generate new original words and phra…

Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models

2025-10-14 · Ziliang Qiu, Renfen Hu arxiv

The evaluation of LLMs' creativity represents a crucial research domain, though challenges such as data contamination and costly human assessments often impede progress. Drawing inspiration from human creativity assessme…

Large Language Models show both individual and collective creativity comparable to humans

2024-12-04 · Luning Sun, Yuzhuo Yuan, Yuan YAO, Yanyan Li 외

Artificial intelligence has, so far, largely automated routine tasks, but what does it mean for the future of work if Large Language Models (LLMs) show creativity comparable to humans? To measure the creativity of LLMs h…

Putting GPT-3's Creativity to the (Alternative Uses) Test

2022-06-10 · Claire Stevenson, Iris Smal, Matthijs Baas, Raoul Grasman 외

AI large language models have (co-)produced amazing written works from newspaper articles to novels and poetry. These works meet the standards of the standard definition of creativity: being original and useful, and some…

ArticlesLanguage ModelingLanguage ModellingLarge Language Model

CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity

2025-10-23 · Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel 외 arxiv

Creativity is often seen as a hallmark of human intelligence. While large language models(LLMs) are increasingly perceived as generating creative text, there is still no cross-domain and scalable framework to evaluate th…

Logical Reasoning