paper-with-me

홈 › Papers

IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation

2026-07-11 · Hengquan Guo arxiv

Scientific ideation unfolds over multiple stages, including literature search, paper reading, tool use, claim checking, cross-paper synthesis, brainstorming, rejection of weak directions, and iterative writing. Yet most existing resources capture isolated components or final artifacts rather than the process connecting them. We introduce IdeaTrail, a dataset of 1,170 multi-turn trajectories for scientific ideation and proposal generation. Each trajectory follows a research process from evidence gathering to either idea selection or proposal construction, jointly recording tool use, acquired evidence, intermediate artifacts, and reasoning. IdeaTrail is synthesized from human-selected research papers and proposal artifacts through a Generator--Advisor loop. The Generator produces the visible sequence of actions, observations, and artifact edits, while the Advisor uses the full generation context to check grounding, causal order, naturalness, and leakage from hidden targets. This reverse-to-forward design keeps trajectories aligned with real scientific artifacts while retaining the uncertainty, evidence use, and staged convergence characteristic of research practice. IdeaTrail provides both reusable process supervision and a general recipe for constructing scientific-research-agent data.

📄 PDF Abstract BibTeX arXiv:2607.10144

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

2026-08-14 · Yanlin Fei, Nazhou Liu, Xinmiao Yu, Shaolong Chen 외 hf

AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final…

FrontierChallenge: Evaluating Scientific Workflow Completion

2026-08-25 · Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang 외 hf

Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domai…

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

2026-02-13 · Yujiong Shen, Yajie Yang, Zhiheng Xi, Binze Hu 외 arxiv

Scientific reasoning inherently demands integrating sophisticated toolkits to navigate domain-specific knowledge. Yet, current benchmarks largely overlook agents' ability to orchestrate tools for such rigorous workflows.…

The (R)evolution of Scientific Workflows in the Agentic AI Era: Towards Autonomous Science

2025-09-12 · Woong Shin, Renan Souza, Daniel Rosendo, Frédéric Suter 외 arxiv

Modern scientific discovery increasingly requires coordinating distributed facilities and heterogeneous resources, forcing researchers to act as manual workflow coordinators rather than scientists. Advances in AI leading…

What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity

2025-11-19 · Alexis Audran-Reiss, Jordi Armengol-Estapé, Karen Hambardzumyan, Amar Budhiraja 외 arxiv

AI research agents offer the promise to accelerate scientific progress by automating the design, implementation, and training of machine learning models. However, the field is still in its infancy, and the key factors dr…