paper-with-me

홈 › Papers

Planning in Strawberry Fields: Evaluating and Improving the Planning and Scheduling Capabilities of LRM o1

2024-10-03 · Karthik Valmeekam, Kaya Stechly, Atharva Gundawar, Subbarao Kambhampati

The ability to plan a course of action that achieves a desired state of affairs has long been considered a core competence of intelligent agents and has been an integral part of AI research since its inception. With the advent of large language models (LLMs), there has been considerable interest in the question of whether or not they possess such planning abilities, but -- despite the slew of new private and open source LLMs since GPT3 -- progress has remained slow. OpenAI claims that their recent o1 (Strawberry) model has been specifically constructed and trained to escape the normal limitations of autoregressive LLMs -- making it a new kind of model: a Large Reasoning Model (LRM). In this paper, we evaluate the planning capabilities of two LRMs (o1-preview and o1-mini) on both planning and scheduling benchmarks. We see that while o1 does seem to offer significant improvements over autoregressive LLMs, this comes at a steep inference cost, while still failing to provide any guarantees over what it generates. We also show that combining o1 models with external verifiers -- in a so-called LRM-Modulo system -- guarantees the correctness of the combined system's output while further improving performance.

📄 PDF Abstract BibTeX arXiv:2410.02162

Code (2)

karthikv792/gpt-plan-benchmark
karthikv792/llms-planning

Tasks

Scheduling

Similar Papers 제목 키워드 기반

AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems

2026-01-16 · Weiyi Wang, Xinchi Chen, Jingjing Gong, Xuanjing Huang 외 arxiv

Recent advances in agentic Large Language Models (LLMs) have positioned them as generalist planners capable of reasoning and acting across diverse tasks. However, existing agent benchmarks largely focus on symbolic or we…

NATURAL PLAN: Benchmarking LLMs on Natural Language Planning

2024-06-06 · Huaixiu Steven Zheng, Swaroop Mishra, Hugh Zhang, Xinyun Chen 외

We introduce NATURAL PLAN, a realistic planning benchmark in natural language containing 3 key tasks: Trip Planning, Meeting Planning, and Calendar Scheduling. We focus our evaluation on the planning capabilities of LLMs…

BenchmarkingScheduling

PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities

2025-04-21 · Haoming Li, Zhaoliang Chen, Jonathan Zhang, Fei Liu

Planning is central to agents and agentic AI. The ability to plan, e.g., creating travel itineraries within a budget, holds immense potential in both scientific and commercial contexts. Moreover, optimal plans tend to re…

Scheduling

TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks

2025-11-03 · Hanwen Xu, Xuyao Huang, Yuzhe Liu, Kai Yu 외 arxiv

Large language model (LLM) agents have exhibited strong problem-solving competence across domains like research and coding. Yet, it remains underexplored whether LLM agents can tackle compounding real-world problems that…

Reinforcement Learning

LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

2024-09-20 · Karthik Valmeekam, Kaya Stechly, Subbarao Kambhampati

The ability to plan a course of action that achieves a desired state of affairs has long been considered a core competence of intelligent agents and has been an integral part of AI research since its inception. With the …