paper-with-me

홈 › Papers

Beyond Winning and Losing: Modeling Human Motivations and Behaviors Using Inverse Reinforcement Learning

2018-07-01 · Baoxiang Wang, Tongfang Sun, Xianjun Sam Zheng

In recent years, reinforcement learning (RL) methods have been applied to model gameplay with great success, achieving super-human performance in various environments, such as Atari, Go, and Poker. However, those studies mostly focus on winning the game and have largely ignored the rich and complex human motivations, which are essential for understanding different players' diverse behaviors. In this paper, we present a novel method called Multi-Motivation Behavior Modeling (MMBM) that takes the multifaceted human motivations into consideration and models the underlying value structure of the players using inverse RL. Our approach does not require the access to the dynamic of the system, making it feasible to model complex interactive environments such as massively multiplayer online games. MMBM is tested on the World of Warcraft Avatar History dataset, which recorded over 70,000 users' gameplay spanning three years period. Our model reveals the significant difference of value structures among different player groups. Using the results of motivation modeling, we also predict and explain their diverse gameplay behaviors and provide a quantitative assessment of how the redesign of the game environment impacts players' behaviors.

📄 PDF Abstract BibTeX arXiv:1807.00366

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Beyond Winning and Losing: Modeling Human Motivations and Behaviors with Vector-valued Inverse Reinforcement Learning

2018-09-27 · Baoxiang Wang, Tongfang Sun, Xianjun Sam Zheng

In recent years, reinforcement learning methods have been applied to model gameplay with great success, achieving super-human performance in various environments, such as Atari, Go and Poker. However, those studies mostl…

Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization

2024-08-14 · Yuxin Jiang, Bo Huang, YuFei Wang, Xingshan Zeng 외

Direct preference optimization (DPO), a widely adopted offline preference optimization algorithm, aims to align large language models (LLMs) with human-desired behaviors using pairwise preference data. However, the winni…

InformativenessInstruction FollowingMath

COMMA: Modeling Relationship among Motivations, Emotions and Actions in Language-based Human Activities

2022-09-14 · COLING 2022 10 · Yuqiang Xie, Yue Hu, Wei Peng, Guanqun Bi 외

Motivations, emotions, and actions are inter-related essential factors in human activities. While motivations and emotions have long been considered at the core of exploring how people take actions in human activities, t…

Action Generation

Sounding Like a Winner? Prosodic Differences in Post-Match Interviews

2025-06-02 · Sofoklis Kakouros, Haoyu Chen

This study examines the prosodic characteristics associated with winning and losing in post-match tennis interviews. Additionally, this research explores the potential to classify match outcomes solely based on post-matc…

Self-Supervised Learning

New fairness criteria for truncated ballots in multi-winner ranked-choice elections

2024-08-07 · Adam Graham-Squire, Matthew I. Jones, David McCune

In real-world elections where voters cast preference ballots, voters often provide only a partial ranking of the candidates. Despite this empirical reality, prior social choice literature frequently analyzes fairness cri…

Fairness