paper-with-me

Papers

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models

2025-05-13 · Xiaoyang Chen, Xinan Dai, Yu Du, Qian Feng, Naixu Guo, Tingshuo Gu, Yuting Gao, Yingyi Gao, Xudong Han, Xiang Jiang, Yilin Jin, Hongyi Lin, Shisheng Lin, Xiangnan Li, Yuante Li, Yixing Li, Zhentao Lai, Zilu Ma, Yingrong Peng, Jiacheng Qian, Hao-Yu Sun, Jianbo Sun, ZiRui Wang, Siwei Wu, Zian Wang, Bin Xu, Jianghao Xu, Yiyang Yu, Zichuan Yang, Hongji Zha, Ruichong Zhang

To advance the mathematical proficiency of large language models (LLMs), the DeepMath team has launched an open-source initiative aimed at developing an open mathematical LLM and systematically evaluating its mathematical creativity. This paper represents the initial contribution of this initiative. While recent developments in mathematical LLMs have predominantly emphasized reasoning skills, as evidenced by benchmarks on elementary to undergraduate-level mathematical tasks, the creative capabilities of these models have received comparatively little attention, and evaluation datasets remain scarce. To address this gap, we propose an evaluation criteria for mathematical creativity and introduce DeepMath-Creative, a novel, high-quality benchmark comprising constructive problems across algebra, geometry, analysis, and other domains. We conduct a systematic evaluation of mainstream LLMs' creative problem-solving abilities using this dataset. Experimental results show that even under lenient scoring criteria -- emphasizing core solution components and disregarding minor inaccuracies, such as small logical gaps, incomplete justifications, or redundant explanations -- the best-performing model, O3 Mini, achieves merely 70% accuracy, primarily on basic undergraduate-level constructive tasks. Performance declines sharply on more complex problems, with models failing to provide substantive strategies for open problems. These findings suggest that, although current LLMs display a degree of constructive proficiency on familiar and lower-difficulty problems, such performance is likely attributable to the recombination of memorized patterns rather than authentic creative insight or novel synthesis.

📄 PDF Abstract BibTeX arXiv:2505.08744

Code (1)

deepmathllm/deepmath 공식 구현

Similar Papers 제목 키워드 기반

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

2025-04-15 · Zhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu 외

The capacity for complex mathematical reasoning is a key benchmark for artificial intelligence. While reinforcement learning (RL) applied to LLMs shows promise, progress is significantly hindered by the lack of large-sca…

Mathematical ReasoningReinforcement Learning (RL)

Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems

2024-10-24 · Junyi Ye, Jingyi Gu, Xinyun Zhao, Wenpeng Yin 외

The mathematical capabilities of AI systems are complex and multifaceted. Most existing research has predominantly focused on the correctness of AI-generated solutions to mathematical problems. In this work, we argue tha…

Mathematical Reasoning

Evaluating Creativity in Computational Co-Creative Systems

2018-07-25 · Pegah Karimi, Kazjon Grace, Mary Lou Maher, Nicholas Davis

This paper provides a framework for evaluating creativity in co-creative systems: those that involve computer programs collaborating with human users on creative tasks. We situate co-creative systems within a broader con…

Creative Invention Benchmark

2018-05-09 · Matthew Guzdial, Nicholas Liao, Vishwa Shah, Mark O. Riedl

In this paper we present the Creative Invention Benchmark (CrIB), a 2000-problem benchmark for evaluating a particular facet of computational creativity. Specifically, we address combinational p-creativity, the creativit…

Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment

2025-10-01 · Vanya Bannihatti Kumar, Divyanshu Goyal, Akhil Eppa, Neel Bhandari arxiv

Modern large language models (LLMs) excel at objective tasks such as evaluating mathematical reasoning and factual accuracy, yet they falter when faced with the nuanced, subjective nature of assessing creativity. In this…

Mathematical Reasoning