paper-with-me

홈 › Papers

PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments

2025-09-07 · Olivier Schipper, Yudi Zhang, Yali Du, Mykola Pechenizkiy, Meng Fang arxiv

LLM-based agents have shown promise in various cooperative and strategic reasoning tasks, but their effectiveness in competitive multi-agent environments remains underexplored. To address this gap, we introduce PillagerBench, a novel framework for evaluating multi-agent systems in real-time competitive team-vs-team scenarios in Minecraft. It provides an extensible API, multi-round testing, and rule-based built-in opponents for fair, reproducible comparisons. We also propose TactiCrafter, an LLM-based multi-agent system that facilitates teamwork through human-readable tactics, learns causal dependencies, and adapts to opponent strategies. Our evaluation demonstrates that TactiCrafter outperforms baseline approaches and showcases adaptive learning through self-play. Additionally, we analyze its learning process and strategic evolution over multiple game episodes. To encourage further research, we have open-sourced PillagerBench, fostering advancements in multi-agent AI for competitive environments.

📄 PDF Abstract BibTeX arXiv:2509.06235

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enabling Rapid Shared Human-AI Mental Model Alignment via the After-Action Review

2025-03-25 · Edward Gu, Ho Chit Siu, Melanie Platt, Isabelle Hurley 외

In this work, we present two novel contributions toward improving research in human-machine teaming (HMT): 1) a Minecraft testbed to accelerate testing and deployment of collaborative AI agents and 2) a tool to allow use…

Minecraft

TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft

2024-12-06 · Qian Long, Zhi Li, Ran Gong, Ying Nian Wu 외

Collaboration is a cornerstone of society. In the real world, human teammates make use of multi-sensory data to tackle challenging tasks in ever-changing environments. It is essential for embodied agents collaborating in…

Imitation LearningMinecraft

Probabilistic Modeling of Human Teams to Infer False Beliefs

2023-10-19 · Paulo Soares, Adarsh Pyarelal, Kobus Barnard

We develop a probabilistic graphical model (PGM) for artificially intelligent (AI) agents to infer human beliefs during a simulated urban search and rescue (USAR) scenario executed in a Minecraft environment with a team …

AI AgentMinecraft

Evaluation of Human-AI Teams for Learned and Rule-Based Agents in Hanabi

2021-07-15 · NeurIPS 2021 12 · Ho Chit Siu, Jaime D. Pena, Edenna Chen, Yutai Zhou 외

Deep reinforcement learning has generated superhuman AI in competitive games such as Go and StarCraft. Can similar learning techniques create a superior AI teammate for human-machine collaborative games? Will humans pref…

BenchmarkingDeep Reinforcement Learningreinforcement-learningReinforcement Learning+2

BEDD: The MineRL BASALT Evaluation and Demonstrations Dataset for Training and Benchmarking Agents that Solve Fuzzy Tasks

2023-12-05 · NeurIPS 2023 11 · Stephanie Milani, Anssi Kanervisto, Karolis Ramanauskas, Sander Schulhoff 외

The MineRL BASALT competition has served to catalyze advances in learning from human feedback through four hard-to-specify tasks in Minecraft, such as create and photograph a waterfall. Given the completion of two years …

BenchmarkingMinecraft