paper-with-me

Papers

AMONGAGENTS: Evaluating Large Language Models in the Interactive Text-Based Social Deduction Game

2024-07-23 · Yizhou Chi, Lingjun Mao, Zineng Tang

Strategic social deduction games serve as valuable testbeds for evaluating the understanding and inference skills of language models, offering crucial insights into social science, artificial intelligence, and strategic gaming. This paper focuses on creating proxies of human behavior in simulated environments, with Among Us utilized as a tool for studying simulated human behavior. The study introduces a text-based game environment, named AmongAgents, that mirrors the dynamics of Among Us. Players act as crew members aboard a spaceship, tasked with identifying impostors who are sabotaging the ship and eliminating the crew. Within this environment, the behavior of simulated language agents is analyzed. The experiments involve diverse game sequences featuring different configurations of Crewmates and Impostor personality archetypes. Our work demonstrates that state-of-the-art large language models (LLMs) can effectively grasp the game rules and make decisions based on the current context. This work aims to promote further exploration of LLMs in goal-oriented games with incomplete information and complex action spaces, as these settings offer valuable opportunities to assess language model performance in socially driven scenarios.

📄 PDF Abstract BibTeX arXiv:2407.16521

Code (1)

cyzus/among-agents 공식 구현

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria

2023-09-24 · Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim 외

By simply composing prompts, developers can prototype novel generative applications with Large Language Models (LLMs). To refine prototypes into products, however, developers must iteratively revise prompts by evaluating…

Language ModelingLanguage ModellingLarge Language Model

StyleBench: Evaluating Speech Language Models on Conversational Speaking Style Control

2026-03-08 · Haishu Zhao, Aokai Hao, Yuan Ge, Zhenqiang Hong 외 arxiv

Speech language models (SLMs) have significantly extended the interactive capability of text-based Large Language Models (LLMs) by incorporating paralinguistic information. For more realistic interactive experience with …

A Survey on Complex Tasks for Goal-Directed Interactive Agents

2024-09-27 · Mareike Hartmann, Alexander Koller

Goal-directed interactive agents, which autonomously complete tasks through interactions with their environment, can assist humans in various domains of their daily lives. Recent advances in large language models (LLMs) …

LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts

2025-09-20 · Junhao Chen, Jingbo Sun, Xiang Li, Haidong Xin 외 arxiv

As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important. To fully assess LLM intelligence, it is crucial to examine their i…

Evaluating Cognitive Age Alignment in Interactive AI Agents

2026-05-18 · Yifan Shen, Jiawen Zhang, Jian Xu, Junho Kim 외 arxiv

While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across domains ranging from daily life to advanced scientific research, a profo…

Visual Reasoning