paper-with-me

Papers

Evaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines

2025-12-24 · Devesh Saraogi, Rohit Singhee, Dhruv Kumar arxiv

The integration of Large Language Models (LLMs) into the scientific ecosystem raises fundamental questions about the creativity and originality of AI-generated research. Recent work has identified ``smart plagiarism'' as a concern in single-step prompting approaches, where models reproduce existing ideas with terminological shifts. This paper investigates whether agentic workflows -- multi-step systems employing iterative reasoning, evolutionary search, and recursive decomposition -- can generate more novel and feasible research plans. We benchmark five reasoning architectures: Reflection-based iterative refinement, Sakana AI v2 evolutionary algorithms, Google Co-Scientist multi-agent framework, GPT Deep Research (GPT-5.1) recursive decomposition, and Gemini~3 Pro multimodal long-context pipeline. Using evaluations from thirty proposals each on novelty, feasibility, and impact, we find that decomposition-based and long-context workflows achieve mean novelty of 4.17/5, while reflection-based approaches score significantly lower (2.33/5). Results reveal varied performance across research domains, with high-performing workflows maintaining feasibility without sacrificing creativity. These findings support the view that carefully designed multi-stage agentic workflows can advance AI-assisted research ideation.

📄 PDF Abstract BibTeX arXiv:2601.09714

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating $n$-Gram Novelty of Language Models Using Rusty-DAWG

2024-06-18 · William Merrill, Noah A. Smith, Yanai Elazar

How novel are texts generated by language models (LMs) relative to their training corpora? In this work, we investigate the extent to which modern LMs generate $n$-grams from their training data, evaluating both (i) the …

Plans for Evaluating Structured Generative Search Summaries

2026-05-26 · Tetsuya Sakai, Jina Lee, Hanpei Fang, Young-In Song arxiv

We propose a framework for evaluating structured generative search summaries that are placed atop organic web search results. A structured summary, generated by a large language model, typically consists of an overview, …

Evolving and Executing Research Plans via Double-Loop Multi-Agent Collaboration

2025-10-08 · Zhi Zhang, Yan Liu, Zhejing Hu, Gong Chen 외 arxiv

Automating the end-to-end scientific research process poses a fundamental challenge: it requires both evolving high-level plans that are novel and sound, and executing these plans correctly amidst dynamic and uncertain c…

OpenKBP-Opt: An international and reproducible evaluation of 76 knowledge-based planning pipelines

2022-02-16 · Aaron Babier, Rafid Mahmood, Binghao Zhang, Victor G. L. Alves 외

We establish an open framework for developing plan optimization models for knowledge-based planning (KBP) in radiotherapy. Our framework includes reference plans for 100 patients with head-and-neck cancer and high-qualit…

HindSight: Evaluating LLM-Generated Research Ideas via Future Impact

2026-03-16 · Bo Jiang arxiv

Evaluating AI-generated research ideas typically relies on LLM judges or human panels -- both subjective and disconnected from actual research impact. We introduce HindSight, a time-split evaluation framework that measur…