paper-with-me

홈 › Papers

Towards Game-Playing AI Benchmarks via Performance Reporting Standards

2020-07-06 · Vanessa Volz, Boris Naujoks

While games have been used extensively as milestones to evaluate game-playing AI, there exists no standardised framework for reporting the obtained observations. As a result, it remains difficult to draw general conclusions about the strengths and weaknesses of different game-playing AI algorithms. In this paper, we propose reporting guidelines for AI game-playing performance that, if followed, provide information suitable for unbiased comparisons between different AI approaches. The vision we describe is to build benchmarks and competitions based on such guidelines in order to be able to draw more general conclusions about the behaviour of different AI algorithms, as well as the types of challenges different games pose.

📄 PDF Abstract BibTeX arXiv:2007.02742

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Optimal Private Payoff Manipulation against Commitment in Extensive-form Games

2022-06-27 · Yurong Chen, Xiaotie Deng, Yuhao Li

To take advantage of strategy commitment, a useful tactic of playing games, a leader must learn enough information about the follower's payoff function. However, this leaves the follower a chance to provide fake informat…

Form

Rinascimento: Optimising Statistical Forward Planning Agents for Playing Splendor

2019-04-03 · Ivan Bravi, Simon Lucas, Diego Perez-Liebana, Jialin Liu

Game-based benchmarks have been playing an essential role in the development of Artificial Intelligence (AI) techniques. Providing diverse challenges is crucial to push research toward innovation and understanding in mod…

Reasoning Capabilities of Large Language Models. Lessons Learned from General Game Playing

2026-02-22 · Maciej Świechowski, Adam Żychowski, Jacek Mańdziuk arxiv

This paper examines the reasoning capabilities of Large Language Models (LLMs) from a novel perspective, focusing on their ability to operate within formally specified, rule-governed environments. We evaluate four LLMs (…

Valet: A Standardized Testbed of Traditional Imperfect-Information Card Games

2026-03-03 · Mark Goadrich, Achille Morenville, Éric Piette arxiv

AI algorithms for imperfect-information games are typically compared using performance metrics on individual games, making it difficult to assess robustness across game choices. Card games are a natural domain for imperf…

Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare

2025-05-12 · Amara Tariq, Rimita Lahiri, Charles Kahn, Imon Banerjee

The intricate and multifaceted nature of vision language model (VLM) development, adaptation, and application necessitates the establishment of clear and standardized reporting protocols, particularly within the high-sta…

Position