paper-with-me

Papers

The Limits of Automatic Evaluation of Creativity in Large Language Models

2026-08-24 · Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi arxiv

Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity. We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, and compare these judgments with automated objective metrics and LLM-as-a-Judge evaluations. Our experiments reveal substantial misalignment between automatic evaluations and human assessments. In particular, LLM-based judges exhibit a systematic preference for AI-generated stories, consistently favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Furthermore, correlation analyses show that widely used automatic metrics exhibit near-zero alignment with human judgments across both human- and AI-generated stories, suggesting that they fail to capture important dimensions of creativity. These findings highlight fundamental limitations in current approaches to the automatic evaluation of creative text and underscore the difficulty of reducing the multidimensional and subjective nature of creativity to computational metrics.

📄 PDF Abstract BibTeX arXiv:2608.23705

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity

2025-10-23 · Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel 외 arxiv

Creativity is often seen as a hallmark of human intelligence. While large language models(LLMs) are increasingly perceived as generating creative text, there is still no cross-domain and scalable framework to evaluate th…

Logical Reasoning

Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations

2026-05-13 · Kyo Gerrits, Rik van Noord, Ana Guerberof Arenas arxiv

This article investigates the performance of automatic evaluation metrics (AEMs) and LLM-as-a-judge evaluation on literary translation across multiple languages, genres, and translation modalities. The aim is to assess h…

Machine Translation

Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation

2026-04-20 · Ran Zhang, Steffen Eger, Arda Tezcan, Wei Zhao 외 arxiv

Large language models (LLMs) are increasingly used for creative tasks such as literary translation. Yet translational creativity remains underexplored and is rarely evaluated at scale, while source-text comprehension is …

The creative psychometric item generator: a framework for item generation and validation using large language models

2024-08-30 · Antonio Laverghetta Jr., Simone Luchini, Averie Linell, Roni Reiter-Palmon 외

Increasingly, large language models (LLMs) are being used to automate workplace processes requiring a high degree of creativity. While much prior work has examined the creativity of LLMs, there has been little research o…

valid

Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey

2025-11-05 · Fatemeh Shahhosseini, Arash Marioriyad, Ali Momen, Mahdieh Soleymani Baghshah 외 arxiv

Scientific idea generation is central to discovery, requiring the joint satisfaction of novelty and scientific soundness. Unlike standard reasoning or general creative generation, scientific ideation is inherently open-e…