paper-with-me

홈 › Papers

AI Idea Bench 2025: AI Research Idea Generation Benchmark

2025-04-19 · Yansheng Qiu, Haoquan Zhang, Zhaopan Xu, Ming Li, Diping Song, Zheng Wang, Kaipeng Zhang

Large-scale Language Models (LLMs) have revolutionized human-AI interaction and achieved significant success in the generation of novel ideas. However, current assessments of idea generation overlook crucial factors such as knowledge leakage in LLMs, the absence of open-ended benchmarks with grounded truth, and the limited scope of feasibility analysis constrained by prompt design. These limitations hinder the potential of uncovering groundbreaking research ideas. In this paper, we present AI Idea Bench 2025, a framework designed to quantitatively evaluate and compare the ideas generated by LLMs within the domain of AI research from diverse perspectives. The framework comprises a comprehensive dataset of 3,495 AI papers and their associated inspired works, along with a robust evaluation methodology. This evaluation system gauges idea quality in two dimensions: alignment with the ground-truth content of the original papers and judgment based on general reference material. AI Idea Bench 2025's benchmarking system stands to be an invaluable resource for assessing and comparing idea-generation techniques, thereby facilitating the automation of scientific discovery.

📄 PDF Abstract BibTeX arXiv:2504.14191

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarkingscientific discovery

Similar Papers 제목 키워드 기반

LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

2026-08-13 · Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li 외 arxiv

With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for researc…

IdeaBench: Benchmarking Large Language Models for Research Idea Generation

2024-10-31 · Sikun Guo, Amir Hassan Shariatmadari, Guangzhi Xiong, Albert Huang 외

Large Language Models (LLMs) have transformed how people interact with artificial intelligence (AI) systems, achieving state-of-the-art results in various tasks, including scientific discovery and hypothesis generation. …

Benchmarkingscientific discovery

AutoResearch: Insight In, Hallucination Out

2026-08-18 · Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang 외 arxiv

Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two…

Cross-Modal Retrieval

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

2026-07-09 · Yifan Zhou, Qihao Yang, Yan Li, Donggang Li 외 arxiv

Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI…

Predicting Empirical AI Research Outcomes with Language Models

2025-06-01 · Jiaxin Wen, Chenglei Si, Yueh-han Chen, He He 외

Many promising-looking ideas in AI research fail to deliver, but their validation takes substantial human labor and compute. Predicting an idea's chance of success is thus crucial for accelerating empirical AI research, …

Retrieval