paper-with-me

홈 › Papers

Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments

2026-01-12 · Bingyang Ye, Shan Chen, Jingxuan Tu, Chen Liu, Zidi Xiong, Samuel Schmidgall, Danielle S. Bitterman arxiv

Large language models are increasingly being used to assess and forecast research ideas, yet we lack scalable ways to evaluate the quality of models' judgments about these scientific ideas. Towards this goal, we introduce PoT, a semi-verifiable benchmarking framework that links scientific idea judgments to downstream signals that become observable later (e.g., citations and shifts in researchers' agendas). PoT freezes a pre-cutoff snapshot of evidence in an offline sandbox and asks models to forecast post-cutoff outcomes, enabling verifiable evaluation when ground truth arrives, scalable benchmarking without exhaustive expert annotation, and analysis of human-model misalignment against signals such as peer-review awards. In addition, PoT provides a controlled testbed for agent-based research judgments that evaluate scientific ideas, comparing tool-using agents to non-agent baselines under prompt ablations and budget scaling. Across 30,000+ instances spanning four benchmark domains, we find that, compared with non-agent baselines, higher interaction budgets generally improve agent performance, while the benefit of tool use is strongly task-dependent. By combining time-partitioned, future-verifiable targets with an offline sandbox for tool use, PoT supports scalable evaluation of agents on future-facing scientific idea judgment tasks.

📄 PDF Abstract BibTeX arXiv:2601.07606

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context

2024-12-23 · Kai Ruan, Xuan Wang, Jixiang Hong, Peng Wang 외

While Large Language Models (LLMs) have demonstrated remarkable capabilities in scientific tasks, existing evaluation frameworks primarily assess their performance using rich contextual inputs, overlooking their ability …

An Evaluation-Centric Paradigm for Scientific Visualization Agents

2025-09-18 · Kuangshi Ai, Haichao Miao, Zhimin Li, Chaoli Wang 외 arxiv

Recent advances in multi-modal large language models (MLLMs) have enabled increasingly sophisticated autonomous visualization agents capable of translating user intentions into data visualizations. However, measuring pro…

Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation

2026-04-06 · Srujan P Mule, Aniketh Garikaparthi, Manasi Patwardhan arxiv

As language models accelerate scientific research by automating hypothesis generation and implementation, a new bottleneck emerges: evaluating and filtering hundreds of AI-generated ideas without exhaustive experimentati…

Reinforcement Learning

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs

2025-12-23 · Saeed Mohammadzadeh, Erfan Hamdi, Joel Shor, Emma Lejeune arxiv

As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to generate scientifically valid physical models has become a critical gap. Computati…

MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research

2025-05-26 · Hui Chen, Miao Xiong, Yujie Lu, Wei Han 외

Recent advancements in AI agents have demonstrated their growing potential to drive and support scientific discovery. In this work, we introduce MLR-Bench, a comprehensive benchmark for evaluating AI agents on open-ended…

scientific discovery