paper-with-me

홈 › Papers

Who is a Better Player: LLM against LLM

2025-08-05 · Yingjie Zhou, Jiezhang Cao, Farong Wen, Li Xu, Yanwei Jiang, Jun Jia, Ronghui Li, Xiaohong Liu, Yu Zhou, Xiongkuo Min, Jie Guo, Zicheng Zhang, Guangtao Zhai arxiv

Adversarial board games, as a paradigmatic domain of strategic reasoning and intelligence, have long served as both a popular competitive activity and a benchmark for evaluating artificial intelligence (AI) systems. Building on this foundation, we propose an adversarial benchmarking framework to assess the comprehensive performance of Large Language Models (LLMs) through board games competition, compensating the limitation of data dependency of the mainstream Question-and-Answer (Q&A) based benchmark method. We introduce Qi Town, a specialized evaluation platform that supports 5 widely played games and involves 20 LLM-driven players. The platform employs both the Elo rating system and a novel Performance Loop Graph (PLG) to quantitatively evaluate the technical capabilities of LLMs, while also capturing Positive Sentiment Score (PSS) throughout gameplay to assess mental fitness. The evaluation is structured as a round-robin tournament, enabling systematic comparison across players. Experimental results indicate that, despite technical differences, most LLMs remain optimistic about winning and losing, demonstrating greater adaptability to high-stress adversarial environments than humans. On the other hand, the complex relationship between cyclic wins and losses in PLGs exposes the instability of LLMs' skill play during games, warranting further explanation and exploration.

📄 PDF Abstract BibTeX arXiv:2508.04720

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Better Computer Go Player with Neural Network and Long-term Prediction

2015-11-19 · Yuandong Tian, Yan Zhu

Competing with top human players in the ancient game of Go has been a long-term goal of artificial intelligence. Go's high branching factor makes traditional search techniques ineffective, even on leading-edge hardware, …

4kGame of Go

Equilibrium Refinement for the Age of Machines: The One-Sided Quasi-Perfect Equilibrium

2021-12-01 · NeurIPS 2021 12 · Gabriele Farina, Tuomas Sandholm

In two-player zero-sum extensive-form games, Nash equilibrium prescribes optimal strategies against perfectly rational opponents. However, it does not guarantee rational play in parts of the game tree that can only be re…

GAN-Aimbots: Using Machine Learning for Cheating in First Person Shooters

2022-05-14 · Anssi Kanervisto, Tomi Kinnunen, Ville Hautamäki

Playing games with cheaters is not fun, and in a multi-billion-dollar video game industry with hundreds of millions of players, game developers aim to improve the security and, consequently, the user experience of their …

BIG-bench Machine Learning

Strategizing against No-regret Learners

2019-09-30 · NeurIPS 2019 12 · Yuan Deng, Jon Schneider, Balusubramanian Sivan

How should a player who repeatedly plays a game against a no-regret learner strategize to maximize his utility? We study this question and show that under some mild assumptions, the player can always guarantee himself a …

iNNk: A Multi-Player Game to Deceive a Neural Network

2020-07-17 · Jennifer Villareale, Ana Acosta-Ruiz, Samuel Arcaro, Thomas Fox 외

This paper presents iNNK, a multiplayer drawing game where human players team up against an NN. The players need to successfully communicate a secret code word to each other through drawings, without being deciphered by …