paper-with-me

홈 › Papers

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

2026-08-22 · Kun Chen, Haorong Hong, Peizhong Gao, Jianfeng Lin, Tongxu Luo, Yuxuan Xie, Chenxu Liu, Jieling He, Zhongyuan Liu, Zeno Zeng hf

Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability must function together in one executable artifact. Measuring this capability therefore requires evaluation of both game product and the development process. Existing benchmarks often assess the game development capabilities of LLMs by evaluating the final artifact or an isolated development stage. Our analysis of complete human-agent development trajectories identifies three stages that together span the lifecycle of game development with a coding agent: initial game generation, bug diagnosis and repair, and optimization over multiple turns. Therefore, we introduce GameXpert-Bench, which operationalizes the three lifecycle stages as three complementary benchmark tracks. GameGen evaluates complete game creation from a single request in an empty workspace. GameFix evaluates diagnosis and repair when defects are reported or left for the agent to discover. GameOpt evaluates cumulative optimization through request chains seeded by real development trajectories between users and agents. We evaluate each track using live game interaction, deterministic behavioral tests, or final product criteria with regression checks. The suite contains 97 generation tasks across 11 genres; 100 repair tasks from 50 game levels verified by humans, each with 19-27 injected bugs; and 17 optimization chains with six turns and 102 requests. Across the three tracks, current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.

📄 PDF Abstract BibTeX arXiv:2608.21833

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games

2026-05-17 · Wenyu Zhang, Guoliang You, Tianlun, Haotian Zhao 외 arxiv

Coding agents are increasingly used as application builders, yet many evaluations still focus on source code, repository-level tests, or intermediate traces rather than the delivered application. We introduce WebGameBenc…

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

2026-06-16 · Tongxu Luo, Rongsheng Wang, Jiaxi Bi, Chenming Xu 외 arxiv

Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place w…

ISO-Bench: Can Coding Agents Optimize Real-World Inference Workloads?

2026-02-23 · Ayush Nangia, Shikhar Mishra, Aman Gokrani, Paras Chopra arxiv

We introduce ISO-Bench, a benchmark for coding agents to test their capabilities on real-world inference optimization tasks. These tasks were taken from vLLM and SGLang, two of the most popular LLM serving frameworks. Ea…

GameEngineBench: Evaluating Coding Agents on Real C++ Runtime Environments

2026-07-03 · Brian La, Sejoon Chang, Ben Kim, Junyoung Bae 외 arxiv

Game engines provide real-time simulation, rendering, physics, interaction, networking, and asset pipelines, making them valuable not only for games but also for 3D applications in healthcare, robotics, architecture, man…

GameDevBench: Evaluating Agentic Capabilities Through Game Development

2026-02-11 · Wayne Chi, Yixiong Fang, Arnav Yayavaram, Siddharth Yayavaram 외 arxiv

Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of evaluation testbeds that combine the complexity of software development with the ne…