paper-with-me

Papers

Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction

2024-07-18 · Suma Bailis, Jane Friedhoff, Feiyang Chen

This paper introduces Werewolf Arena, a novel framework for evaluating large language models (LLMs) through the lens of the classic social deduction game, Werewolf. In Werewolf Arena, LLMs compete against each other, navigating the game's complex dynamics of deception, deduction, and persuasion. The framework introduces a dynamic turn-taking system based on bidding, mirroring real-world discussions where individuals strategically choose when to speak. We demonstrate the framework's utility through an arena-style tournament featuring Gemini and GPT models. Our results reveal distinct strengths and weaknesses in the models' strategic reasoning and communication. These findings highlight Werewolf Arena's potential as a challenging and scalable LLM benchmark.

📄 PDF Abstract BibTeX arXiv:2407.13943

Code (1)

google/werewolf_arena 공식 구현

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Weight Decay 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Optimal Strategy in Werewolf Game: A Game Theoretic Perspective

2024-08-30 · ST Wang

Werewolf game, also known as Mafia game, is a social deduction game that models the conflict between an informed minority (werewolf group) and an uninformed majority (citizen group). This paper explores the optimal strat…

Werewolf: A Straightforward Game Framework with TTS for Improved User Engagement

2025-05-30 · Qihui Fan, Enfu Nan, Wenbo Li, Lei Lu 외

The growing popularity of social deduction game systems for both business applications and AI research has greatly benefited from the rapid advancements in Large Language Models (LLMs), which now demonstrate stronger rea…

text-to-speechText to Speech

Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies

2025-10-13 · Zirui Song, Yuan Huang, Junchang Liu, Haozhe Luo 외 arxiv

Social deduction games like Werewolf combine language, reasoning, and strategy, providing a testbed for studying natural language and social intelligence. However, most studies reduce the game to LLM-based self-play, yie…

Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf

2023-09-09 · Yuzhuang Xu, Shuo Wang, Peng Li, Fuwen Luo 외

Communication games, which we refer to as incomplete information games that heavily depend on natural language communication, hold significant research value in fields such as economics, social science, and artificial in…

Retrieval

MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind

2025-04-25 · Zheng Zhang, Nuoqian Xiao, Qi Chai, Deheng Ye 외

Large Language Model (LLM) agents have demonstrated impressive capabilities in social deduction games (SDGs) like Werewolf, where strategic reasoning and social deception are essential. However, current approaches remain…

Large Language ModelMultimodal Reasoning