paper-with-me

Papers

Creativity Benchmark: A benchmark for marketing creativity for large language models

2025-09-05 · Ninad Bhat, Kieran Browne, Pip Bingemann arxiv

We introduce Creativity Benchmark, an evaluation framework for large language models (LLMs) in marketing creativity. The benchmark covers 100 brands (12 categories) and three prompt types (Insights, Ideas, Wild Ideas). Human pairwise preferences from 678 practising creatives over 11,012 anonymised comparisons, analysed with Bradley-Terry models, show tightly clustered performance with no model dominating across brands or prompt types: the top-bottom spread is $Δθ\approx 0.45$, which implies a head-to-head win probability of $0.61$; the highest-rated model beats the lowest only about $61\%$ of the time. We also analyse model diversity using cosine distances to capture intra- and inter-model variation and sensitivity to prompt reframing. Comparing three LLM-as-judge setups with human rankings reveals weak, inconsistent correlations and judge-specific biases, underscoring that automated judges cannot substitute for human evaluation. Conventional creativity tests also transfer only partially to brand-constrained tasks. Overall, the results highlight the need for expert human evaluation and diversity-aware workflows.

📄 PDF Abstract BibTeX arXiv:2509.09702

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enhancing Creativity in Large Language Models through Associative Thinking Strategies

2024-05-09 · Pronita Mehrotra, Aishni Parab, Sumit Gulwani

This paper explores the enhancement of creativity in Large Language Models (LLMs) like vGPT-4 through associative thinking, a cognitive process where creative ideas emerge from linking seemingly unrelated concepts. Assoc…

Marketing

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

2025-11-17 · Kaiwen Xue, Chenglong Li, Zhonghong Ou, Guoxin Zhang 외 arxiv

Human-defined creativity is highly abstract, posing a challenge for multimodal large language models (MLLMs) to comprehend and assess creativity that aligns with human judgments. The absence of an existing benchmark furt…

AGC-Bench: Measuring Artificial General Creativity

2026-07-01 · Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai, Anna Attuch 외 arxiv

Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from general intelligence. Both questions now apply to LLMs, but a unified b…

General Knowledge

Playing with Words, Improving with Rewards: Training Language Models for Creative Association

2026-05-27 · Vijeta Deshpande, Namrata Shivagunde, Sherin Muckatira, Hadrien Glaude 외 arxiv

Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLMs need to be creative. Yet the subjective nature of creativity and t…

Reinforcement Learning

Creative Invention Benchmark

2018-05-09 · Matthew Guzdial, Nicholas Liao, Vishwa Shah, Mark O. Riedl

In this paper we present the Creative Invention Benchmark (CrIB), a 2000-problem benchmark for evaluating a particular facet of computational creativity. Specifically, we address combinational p-creativity, the creativit…