paper-with-me

홈 › Papers

Digital Player: Evaluating Large Language Models based Human-like Agent in Games

2025-02-28 · Jiawei Wang, Kai Wang, Shaojie Lin, Runze Wu, Bihan Xu, Lingeng Jiang, Shiwei Zhao, Renyu Zhu, Haoyu Liu, Zhipeng Hu, Zhong Fan, Le Li, Tangjie Lyu, Changjie Fan

With the rapid advancement of Large Language Models (LLMs), LLM-based autonomous agents have shown the potential to function as digital employees, such as digital analysts, teachers, and programmers. In this paper, we develop an application-level testbed based on the open-source strategy game "Unciv", which has millions of active players, to enable researchers to build a "data flywheel" for studying human-like agents in the "digital players" task. This "Civilization"-like game features expansive decision-making spaces along with rich linguistic interactions such as diplomatic negotiations and acts of deception, posing significant challenges for LLM-based agents in terms of numerical reasoning and long-term planning. Another challenge for "digital players" is to generate human-like responses for social interaction, collaboration, and negotiation with human players. The open-source project can be found at https:/github.com/fuxiAIlab/CivAgent.

📄 PDF Abstract BibTeX arXiv:2502.20807

Code (1)

fuxiailab/civagent 공식 구현

Tasks

Decision Making

Similar Papers 제목 키워드 기반

AI in (and for) Games

2021-05-07 · Kostas Karpouzis, George Tsatiris

This chapter outlines the relation between artificial intelligence (AI) / machine learning (ML) algorithms and digital games. This relation is two-fold: on one hand, AI/ML researchers can generate large, in-the-wild data…

Relation

Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game

2024-06-16 · Prisha Samadarshi, Mariam Mustafa, Anushka Kulkarni, Raven Rothkopf 외

The New York Times Connections game has emerged as a popular and challenging pursuit for word puzzle enthusiasts. We collect 438 Connections games to evaluate the performance of state-of-the-art large language models (LL…

PLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games

2024-04-26 · Qinglin Zhu, Runcong Zhao, Bin Liang, Jinhua Du 외

We introduce WellPlay, a reasoning dataset for multi-agent conversational inference in Murder Mystery Games (MMGs). WellPlay comprises 1,482 inferential questions across 12 games, spanning objectives, reasoning, and rela…

Decision MakingLanguage ModelingLanguage ModellingLarge Language Model+2

How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO

2024-04-22 · Man Tik Ng, Hui Tung Tse, Jen-tse Huang, Jingjing Li 외

The role-play ability of Large Language Models (LLMs) has emerged as a popular research direction. However, existing studies focus on imitating well-known public figures or fictional characters, overlooking the potential…

Evaluating Soccer Player: from Live Camera to Deep Reinforcement Learning

2021-01-13 · Paul Garnier, Théophane Gregoir

Scientifically evaluating soccer players represents a challenging Machine Learning problem. Unfortunately, most existing answers have very opaque algorithm training procedures; relevant data are scarcely accessible and a…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)