paper-with-me

Papers

LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context

2024-12-23 · Kai Ruan, Xuan Wang, Jixiang Hong, Peng Wang, Yang Liu, Hao Sun

While Large Language Models (LLMs) have demonstrated remarkable capabilities in scientific tasks, existing evaluation frameworks primarily assess their performance using rich contextual inputs, overlooking their ability to generate novel ideas from minimal information. We introduce LiveIdeaBench, a comprehensive benchmark that evaluates LLMs' scientific creativity and divergent thinking capabilities using single-keyword prompts. Drawing from Guilford's creativity theory, our framework employs a dynamic panel of state-of-the-art LLMs to assess generated ideas across four key dimensions: originality, feasibility, fluency, and flexibility. Through extensive experimentation with 20 leading models across 1,180 keywords spanning 18 scientific domains, we reveal that scientific creative ability shows distinct patterns from general intelligence metrics. Notably, our results demonstrate that models like QwQ-32B-preview achieve comparable creative performance to top-tier models like o1-preview, despite significant gaps in their general intelligence scores. These findings highlight the importance of specialized evaluation frameworks for scientific creativity and suggest that the development of creative capabilities in LLMs may follow different trajectories than traditional problem-solving abilities.

📄 PDF Abstract BibTeX arXiv:2412.17596

Code (1)

x66ccff/liveideabench 공식 구현

Similar Papers 제목 키워드 기반

Combinatorial Creativity: A New Frontier in Generalization Abilities

2025-09-25 · Samuel Schapiro, Sumuk Shashidhar, Alexi Gladstone, Jonah Black 외 arxiv

Artificial intelligence (AI) systems, and Large Language Models (LLMs) in particular, are increasingly employed for creative tasks like scientific idea generation, constituting a form of generalization from training data…

LLM Jaggedness Unlocks Scientific Creativity

2026-05-11 · Shray Mathur, J. Anibal Boscoboinik, Esther H. R. Tsai, Kevin G. Yager arxiv

As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities growing unevenly across tasks, domains, and model scales. In this work, we examine…

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers

2026-05-13 · Samuel Schapiro, Alexi Gladstone, Jonah Black, Heng Ji arxiv

Measuring the creativity of large language models (LLMs) is essential for designing methods that can improve creativity and for enhancing our scientific understanding of this ability. To accomplish this, it has become co…

LLMs can realize combinatorial creativity: generating creative ideas via LLMs for scientific research

2024-12-18 · Tianyang Gu, Jingjin Wang, Zhihao Zhang, HaoHong Li

Scientific idea generation has been extensively studied in creativity theory and computational creativity research, providing valuable frameworks for understanding and implementing creative processes. However, recent wor…

Retrieval

Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey

2025-11-05 · Fatemeh Shahhosseini, Arash Marioriyad, Ali Momen, Mahdieh Soleymani Baghshah 외 arxiv

Scientific idea generation is central to discovery, requiring the joint satisfaction of novelty and scientific soundness. Unlike standard reasoning or general creative generation, scientific ideation is inherently open-e…