paper-with-me

홈 › Papers

MinePlanner: A Benchmark for Long-Horizon Planning in Large Minecraft Worlds

2023-12-20 · William Hill, Ireton Liu, Anita de Mello Koch, Damion Harvey, Nishanth Kumar, George Konidaris, Steven James

We propose a new benchmark for planning tasks based on the Minecraft game. Our benchmark contains 45 tasks overall, but also provides support for creating both propositional and numeric instances of new Minecraft tasks automatically. We benchmark numeric and propositional planning systems on these tasks, with results demonstrating that state-of-the-art planners are currently incapable of dealing with many of the challenges advanced by our new benchmark, such as scaling to instances with thousands of objects. Based on these results, we identify areas of improvement for future planners. Our framework is made available at https://github.com/IretonLiu/mine-pddl/.

📄 PDF Abstract BibTeX arXiv:2312.12891

Code (1)

iretonliu/mine-pddl 공식 구현

Tasks

Minecraft

Similar Papers 제목 키워드 기반

When Robots Do the Chores: A Benchmark and Agent for Long-Horizon Household Task Execution

2026-05-14 · Zilin Zhu, Longteng Guo, Yanghong Mei, Bowen Pang 외 arxiv

Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied AI benchmarks that emphasize short-horizon navigation or manipulation…

SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models

2026-01-28 · Sebastiano Monti, Carlo Nicolini, Gianni Pellegrini, Jacopo Staiano 외 arxiv

Although the capabilities of large language models have been increasingly tested on complex reasoning tasks, their long-horizon planning abilities have not yet been extensively investigated. In this work, we provide a sy…

HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds

2025-08-18 · Petr Anokhin, Roman Khalikov, Stefan Rebrikov, Viktor Volkov 외 arxiv

Large language models (LLMs) perform well on step-by-step reasoning benchmarks such as mathematics and code generation, yet their ability to carry out robust long-horizon planning under realistic constraints remains insu…

Spatial ReasoningCode Generation

CookBench: A Long-Horizon Embodied Planning Benchmark for Complex Cooking Scenarios

2025-08-05 · Muzhen Cai, Xiubo Chen, Yining An, Jiaxin Zhang 외 arxiv

Embodied Planning is dedicated to the goal of creating agents capable of executing long-horizon tasks in complex physical worlds. However, existing embodied planning benchmarks frequently feature short-horizon tasks and …

DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints

2026-01-26 · Yinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu 외 arxiv

While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands ge…