paper-with-me

홈 › Papers

LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design

2026-05-14 · Marilyn Zhang, Tianfeng Chen, Fabián Barzuna, Ankita Rathod, Mark E. Whiting arxiv

LLMs are increasingly deployed in autonomous laboratories, under the assumption that their domain priors and reasoning over iterative feedback let them converge on good designs in fewer iterations than feedback-only baselines. Current iterative scientific design benchmarks, however, score only outcome snapshots at fixed horizons. This leaves the learning trajectory unmeasured, even though the trajectory is what captures learning efficiency, where each iteration saved is a real saving in cost and time. Motivated by this, we examine three evaluation choices that change the conclusions one draws about LLM learning efficiency in iterative scientific design: what to measure, what baseline to compare against, and what to ground against. We introduce LEAPBench, Learning Efficiency in Adaptive Processes, a 55-task framework that pairs a best-so-far area under the curve (AUC) trajectory metric with a classical Bayesian-optimization reference and an audit grounded in published literature. Applied to eight contemporary LLMs, switching from final-outcome to trajectory scoring changes the best-model decision on 53% of tasks at matched horizons, and exposes efficiency gains overlooked by outcome-based scoring. LLMs do not outperform a classical Bayesian baseline. On 16 biology tasks where the oracle's reward signal is aligned with configurations from the published-best design, domain-aware prompting leads to LLM choices that match the published-best's approximately 10 percentage points less often than domain-agnostic prompting at iteration 30. The pattern is sharpest on 6 tasks where the literature-typical and published-best configurations diverge, and domain-agnostic prompting matches the published-best more often on all 6. The trajectory metric also doubles as a tractable training target. Offline reinforcement learning with the metric as a reward improves performance on 14 of 21 held-out tasks.

📄 PDF Abstract BibTeX arXiv:2605.15341

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

2026-06-02 · Po-Nien Kung, Linfeng Song, Dawsen Hwang, Jinsung Yoon 외 arxiv

Large Language Models (LLMs) exhibit strong informal mathematical reasoning but struggle to generate mechanically verifiable proofs in formal languages like Lean. We present LEAP, an agentic framework that enables genera…

Mathematical ReasoningInstruction Following

LEAP-VO: Long-term Effective Any Point Tracking for Visual Odometry

2024-01-03 · CVPR 2024 1 · Weirong Chen, Le Chen, Rui Wang, Marc Pollefeys

Visual odometry estimates the motion of a moving camera based on visual input. Existing methods, mostly focusing on two-view point tracking, often ignore the rich temporal context in the image sequence, thereby overlooki…

Point TrackingVisual Odometry

POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization

2025-09-26 · Ziqing Wang, Yibo Wen, William Pattie, Xiao Luo 외 arxiv

Lead optimization in drug discovery requires efficiently navigating vast chemical space through iterative cycles to enhance molecular properties while preserving structural similarity to the original lead compound. Despi…

Reinforcement LearningInstruction FollowingDrug Discovery

Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback

2024-10-07 · Sanjiban Choudhury, Paloma Sodhi

While large language models (LLMs) show impressive decision-making abilities, current methods lack a mechanism for automatic self-improvement from errors during task execution. We propose LEAP, an iterative fine-tuning f…

Decision Makingtext-based games

Planning with Sequence Models through Iterative Energy Minimization

2023-03-28 · Hongyi Chen, Yilun Du, Yiye Chen, Joshua Tenenbaum 외

Recent works have shown that sequence modeling can be effectively used to train reinforcement learning (RL) policies. However, the success of applying existing sequence models to planning, in which we wish to obtain a tr…

Language ModelingLanguage ModellingReinforcement Learning (RL)