paper-with-me

Papers

HindSight: Evaluating LLM-Generated Research Ideas via Future Impact

2026-03-16 · Bo Jiang arxiv

Evaluating AI-generated research ideas typically relies on LLM judges or human panels -- both subjective and disconnected from actual research impact. We introduce HindSight, a time-split evaluation framework that measures idea quality by matching generated ideas against real future publications and scoring them by citation impact and venue acceptance. Using a temporal cutoff~$T$, we restrict an idea generation system to pre-$T$ literature, then evaluate its outputs against papers published in the subsequent 30 months. Experiments across 10 AI/ML research topics reveal a striking disconnect: LLM-as-Judge finds no significant difference between retrieval-augmented and vanilla idea generation ($p{=}0.584$), while HindSight shows the retrieval-augmented system produces 2.5$\times$ higher-scoring ideas ($p{<}0.001$). Moreover, HindSight scores are \emph{negatively} correlated with LLM-judged novelty ($ρ{=}{-}0.29$, $p{<}0.01$), suggesting that LLMs systematically overvalue novel-sounding ideas that never materialize in real research.

📄 PDF Abstract BibTeX arXiv:2603.15164

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can Large Language Models Unlock Novel Scientific Research Ideas?

2024-09-10 · Sandeep Kumar, Tirthankar Ghosal, Vinayak Goyal, Asif Ekbal

"An idea is nothing more nor less than a new combination of old elements" (Young, J.W.). The widespread adoption of Large Language Models (LLMs) and publicly available ChatGPT have marked a significant turning point in t…

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas

2025-06-25 · Chenglei Si, Tatsunori Hashimoto, Diyi Yang

Large Language Models (LLMs) have shown promise in accelerating the scientific research pipeline. A key capability for this process is the ability to generate novel research ideas, and prior studies have found settings i…

AI Research Agents Narrow Scientific Exploration

2026-05-27 · Yixuan Tang, Yi Yang arxiv

AI research agents now support large-scale AI-assisted scientific discovery. We examine whether AI-generated ideas broaden scientific exploration or primarily reinforce existing work. Using five agent frameworks and five…

Understanding EFL Student Idea Generation Strategies for Creative Writing with NLG Tools

2022-06-04 · David James Woo, Yanzhi Wang, Hengky Susanto, Kai Guo

Natural language generation (NLG) is a process within artificial intelligence where computer systems produce human-comprehensible language texts from information. English as a foreign language (EFL) students' use of NLG …

Text Generation

Interesting Scientific Idea Generation using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders

2024-05-27 · Xuemei Gu, Mario Krenn

The rapid growth of scientific literature makes it challenging for researchers to identify novel and impactful ideas, especially across disciplines. Modern artificial intelligence (AI) systems offer new approaches, poten…

Knowledge GraphsLanguage ModelingLanguage ModellingLarge Language Model