paper-with-me

홈 › Papers

Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

2024-09-06 · Chenglei Si, Diyi Yang, Tatsunori Hashimoto

Recent advancements in large language models (LLMs) have sparked optimism about their potential to accelerate scientific discovery, with a growing number of works proposing research agents that autonomously generate and validate new ideas. Despite this, no evaluations have shown that LLM systems can take the very first step of producing novel, expert-level ideas, let alone perform the entire research process. We address this by establishing an experimental design that evaluates research idea generation while controlling for confounders and performs the first head-to-head comparison between expert NLP researchers and an LLM ideation agent. By recruiting over 100 NLP researchers to write novel ideas and blind reviews of both LLM and human ideas, we obtain the first statistically significant conclusion on current LLM capabilities for research ideation: we find LLM-generated ideas are judged as more novel (p < 0.05) than human expert ideas while being judged slightly weaker on feasibility. Studying our agent baselines closely, we identify open problems in building and evaluating research agents, including failures of LLM self-evaluation and their lack of diversity in generation. Finally, we acknowledge that human judgements of novelty can be difficult, even by experts, and propose an end-to-end study design which recruits researchers to execute these ideas into full projects, enabling us to study whether these novelty and feasibility judgements result in meaningful differences in research outcome.

📄 PDF Abstract BibTeX arXiv:2409.04109

Code (3)

NoviScl/AI-Researcher 공식 구현
sharad461/MKA-hallucination pytorch
simplaj/AI-Researcher-Spark

Tasks

Experimental Designscientific discovery

Similar Papers 제목 키워드 기반

Measuring the Gap Between Human and LLM Research Ideas

2026-07-01 · Ziyu Chen, Yilun Zhao, Arman Cohan arxiv

LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference. We instead ask: how far are current LLM-generated ideas from …

Interesting Scientific Idea Generation using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders

2024-05-27 · Xuemei Gu, Mario Krenn

The rapid growth of scientific literature makes it challenging for researchers to identify novel and impactful ideas, especially across disciplines. Modern artificial intelligence (AI) systems offer new approaches, poten…

Knowledge GraphsLanguage ModelingLanguage ModellingLarge Language Model

Can Large Language Models Unlock Novel Scientific Research Ideas?

2024-09-10 · Sandeep Kumar, Tirthankar Ghosal, Vinayak Goyal, Asif Ekbal

"An idea is nothing more nor less than a new combination of old elements" (Young, J.W.). The widespread adoption of Large Language Models (LLMs) and publicly available ChatGPT have marked a significant turning point in t…

Towards Execution-Grounded Automated AI Research

2026-01-20 · Chenglei Si, Zitong Yang, Yejin Choi, Emmanuel Candès 외 arxiv

Automated AI research holds great potential to accelerate scientific discovery. However, current LLMs often generate plausible-looking but ineffective ideas. Execution grounding may help, but it is unclear whether automa…

Reinforcement Learning

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas

2025-06-25 · Chenglei Si, Tatsunori Hashimoto, Diyi Yang

Large Language Models (LLMs) have shown promise in accelerating the scientific research pipeline. A key capability for this process is the ability to generate novel research ideas, and prior studies have found settings i…