paper-with-me

홈 › Papers

Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas

2026-03-11 · Tim Schopf, Michael Färber arxiv

Judging the novelty of research ideas is crucial for advancing science, enabling the identification of unexplored directions, and ensuring contributions meaningfully extend existing knowledge rather than reiterate minor variations. However, given the exponential growth of scientific literature, manually judging the novelty of research ideas through literature reviews is labor-intensive, subjective, and infeasible at scale. Therefore, recent efforts have proposed automated approaches for research idea novelty judgment. Yet, evaluation of these approaches remains largely inconsistent and is typically based on non-standardized human evaluations, hindering large-scale, comparable evaluations. To address this, we introduce RINoBench, the first comprehensive benchmark for large-scale evaluation of research idea novelty judgments. It comprises 1,381 research ideas derived from and judged by human experts as well as nine automated evaluation metrics designed to assess both rubric-based novelty scores and textual justifications of novelty judgments. Using this benchmark, we evaluate several state-of-the-art large language models (LLMs) on their ability to judge the novelty of research ideas. Our findings reveal that while LLM-generated reasoning closely mirrors human rationales, this alignment does not reliably translate into accurate novelty judgments, which diverge significantly from human gold standard judgments - even among leading reasoning-capable models. Data and code available at: https://github.com/TimSchopf/RINoBench.

📄 PDF Abstract BibTeX arXiv:2603.10303

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

2026-08-13 · Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li 외 arxiv

With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for researc…

Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

2026-08-26 · Tim Schopf, Tobias Schreieder, Akiko Aizawa arxiv

Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we invest…

Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments

2026-01-12 · Bingyang Ye, Shan Chen, Jingxuan Tu, Chen Liu 외 arxiv

Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientific ideas. Towards this goal, we introduc…

AI Idea Bench 2025: AI Research Idea Generation Benchmark

2025-04-19 · Yansheng Qiu, Haoquan Zhang, Zhaopan Xu, Ming Li 외

Large-scale Language Models (LLMs) have revolutionized human-AI interaction and achieved significant success in the generation of novel ideas. However, current assessments of idea generation overlook crucial factors such…

Benchmarkingscientific discovery

A Novel Mathematical Framework for Objective Characterization of Ideas

2024-09-11 · B. Sankar, Dibakar Sen

The demand for innovation in product design necessitates a prolific ideation phase. Conversational AI (CAI) systems that use Large Language Models (LLMs) such as GPT (Generative Pre-trained Transformer) have been shown t…

Diversity