paper-with-me

Papers

LLM Jaggedness Unlocks Scientific Creativity

2026-05-11 · Shray Mathur, J. Anibal Boscoboinik, Esther H. R. Tsai, Kevin G. Yager arxiv

As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities growing unevenly across tasks, domains, and model scales. In this work, we examine this dynamic jaggedness through the lens of scientific idea generation. We introduce SciAidanBench, a benchmark of open-ended scientific questions designed to measure the scientific creativity of large language models (LLMs). Given a scientific question, models are asked to generate as many unique and coherent ideas as possible, with the total number of valid responses serving as a proxy for creative potential. Evaluating 19 base models across 8 providers (30 total variants including reasoning versions), we find that jaggedness manifests both across models and within models. First, in a cross-task comparison between general and scientific creativity, improvements in general creativity do not translate uniformly to scientific creativity, revealing divergent capability profiles across models. Second, at the prompt level, stronger models do not improve uniformly; instead, they exhibit high variability, with bursts of creativity on some questions and limited performance on others. Third, at the domain level, individual models display uneven strengths across scientific subfields, reflecting fragmented internal capability profiles. Finally, we show that this jaggedness can be harnessed. We explore mechanisms of inference-time compute, knowledge pooling, and brainstorming to combine models effectively and construct meta-model ensembles that outperform any single model. Our results position jaggedness not as a limitation, but as a resource, a structural feature of AI progress that, when understood and leveraged, can amplify LLM-driven scientific creativity.

📄 PDF Abstract BibTeX arXiv:2605.10574

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers

2026-05-13 · Samuel Schapiro, Alexi Gladstone, Jonah Black, Heng Ji arxiv

Measuring the creativity of large language models (LLMs) is essential for designing methods that can improve creativity and for enhancing our scientific understanding of this ability. To accomplish this, it has become co…

AI and the Decentering of Disciplinary Creativity

2025-10-27 · Eamon Duede arxiv

This paper examines the role of artificial intelligence in scientific problem-solving, with a focus on its implications for disciplinary creativity. Drawing on recent work in the philosophy of creativity, I distinguish b…

Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey

2025-11-05 · Fatemeh Shahhosseini, Arash Marioriyad, Ali Momen, Mahdieh Soleymani Baghshah 외 arxiv

Scientific idea generation is central to discovery, requiring the joint satisfaction of novelty and scientific soundness. Unlike standard reasoning or general creative generation, scientific ideation is inherently open-e…

Transformational Creativity in Science: A Graphical Theory

2025-04-25 · Samuel Schapiro, Jonah Black, Lav R. Varshney

Creative processes are typically divided into three types: combinatorial, exploratory, and transformational. Here, we provide a graphical theory of transformational scientific creativity, synthesizing Boden's insight tha…

LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context

2024-12-23 · Kai Ruan, Xuan Wang, Jixiang Hong, Peng Wang 외

While Large Language Models (LLMs) have demonstrated remarkable capabilities in scientific tasks, existing evaluation frameworks primarily assess their performance using rich contextual inputs, overlooking their ability …