paper-with-me

홈 › Papers

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

2026-01-09 · Dawei Wang, Chengming Zhou, Di Zhao, Xinyuan Liu, Marci Chi Ma, Gary Ushaw, Richard Davison arxiv

Recent breakthroughs in Large Language Models (LLMs) have positioned them as a promising paradigm for agents, with long-term planning and decision-making emerging as core general-purpose capabilities for adapting to diverse scenarios and tasks. Real-time strategy (RTS) games serve as an ideal testbed for evaluating these two capabilities, as their inherent gameplay requires both macro-level strategic planning and micro-level tactical adaptation and action execution. Existing RTS game-based environments either suffer from relatively high computational demands or lack support for textual observations, which has constrained the use of RTS games for LLM evaluation. Motivated by this, we present TowerMind, a novel environment grounded in the tower defense (TD) subgenre of RTS games. TowerMind preserves the key evaluation strengths of RTS games for assessing LLMs, while featuring low computational demands and a multimodal observation space, including pixel-based, textual, and structured game-state representations. In addition, TowerMind supports the evaluation of model hallucination and provides a high degree of customizability. We design five benchmark levels to evaluate several widely used LLMs under different multimodal input settings. The results reveal a clear performance gap between LLMs and human experts across both capability and hallucination dimensions. The experiments further highlight key limitations in LLM behavior, such as inadequate planning validation, a lack of multifinality in decision-making, and inefficient action use. We also evaluate two classic reinforcement learning algorithms: Ape-X DQN and PPO. By offering a lightweight and multimodal design, TowerMind complements the existing RTS game-based environment landscape and introduces a new benchmark for the AI agent field. The source code is publicly available on GitHub(https://github.com/tb6147877/TowerMind).

📄 PDF Abstract BibTeX arXiv:2601.05899

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Obstacle Tower: A Generalization Challenge in Vision, Control, and Planning

2019-02-04 · Arthur Juliani, Ahmed Khalifa, Vincent-Pierre Berges, Jonathan Harper 외

The rapid pace of recent research in AI has been driven in part by the presence of fast and challenging simulation environments. These environments often take the form of games; with tasks ranging from simple board games…

Atari GamesBoard Games

On the recognition of the game type based on physiological signals and eye tracking

2023-10-26 · Łukasz Czekaj, Łukasz Radzinski, Mateusz Kolimaga, Jakub Domaszewicz 외

Automated interpretation of signals yields many impressive applications from the area of affective computing and human activity recognition (HAR). In this paper we ask the question about possibility of cognitive activity…

Activity RecognitionHuman Activity RecognitionPerson Recognition

Towards a Deep Reinforcement Learning Approach for Tower Line Wars

2017-12-17 · Per-Arne Andersen, Morten Goodwin, Ole-Christoffer Granmo

There have been numerous breakthroughs with reinforcement learning in the recent years, perhaps most notably on Deep Reinforcement Learning successfully playing and winning relatively advanced computer games. There is un…

Deep Reinforcement LearningQ-Learningreinforcement-learningReinforcement Learning+3

Learning Cyber Defence Tactics from Scratch with Multi-Agent Reinforcement Learning

2023-08-25 · Jacob Wiebe, Ranwa Al Mallah, Li Li

Recent advancements in deep learning techniques have opened new possibilities for designing solutions for autonomous cyber defence. Teams of intelligent agents in computer network defence roles may reveal promising avenu…

Multi-agent Reinforcement Learningreinforcement-learning

Learning to Communicate in Multi-Agent Reinforcement Learning for Autonomous Cyber Defence

2025-07-19 · Faizan Contractor, Li Li, Ranwa Al Mallah arxiv

Popular methods in cooperative Multi-Agent Reinforcement Learning with partially observable environments typically allow agents to act independently during execution, which may limit the coordinated effect of the trained…

Multi-agent Reinforcement Learning