Beyond Winning and Losing: Modeling Human Motivations and Behaviors with Vector-valued Inverse Reinforcement Learning
In recent years, reinforcement learning methods have been applied to model gameplay with great success, achieving super-human performance in various environments, such as Atari, Go and Poker. However, those studies mostly focus on winning the game and have largely ignored the rich and complex human motivations, which are essential for understanding the agents' diverse behavior. In this paper, we present a multi-motivation behavior modeling which investigates the multifaceted human motivations and models the underlying value structure of the agents. Our approach extends inverse RL to the vectored-valued setting which imposes a much weaker assumption than previous studies. The vectorized rewards incorporate Pareto optimality, which is a powerful tool to explain a wide range of behavior by its optimality. For practical assessment, our algorithm is tested on the World of Warcraft Avatar History dataset spanning three years of the gameplay. Our experiments demonstrate the improvement over the scalarization-based methods on real-world problem settings.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Beyond Winning and Losing: Modeling Human Motivations and Behaviors Using Inverse Reinforcement Learning
In recent years, reinforcement learning (RL) methods have been applied to model gameplay with great success, achieving super-human performance in various environments, such as Atari, Go, and Poker. However, those studies…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
Direct preference optimization (DPO), a widely adopted offline preference optimization algorithm, aims to align large language models (LLMs) with human-desired behaviors using pairwise preference data. However, the winni…
InformativenessInstruction FollowingMathCOMMA: Modeling Relationship among Motivations, Emotions and Actions in Language-based Human Activities
Motivations, emotions, and actions are inter-related essential factors in human activities. While motivations and emotions have long been considered at the core of exploring how people take actions in human activities, t…
Action GenerationSounding Like a Winner? Prosodic Differences in Post-Match Interviews
This study examines the prosodic characteristics associated with winning and losing in post-match tennis interviews. Additionally, this research explores the potential to classify match outcomes solely based on post-matc…
Self-Supervised LearningNew fairness criteria for truncated ballots in multi-winner ranked-choice elections
In real-world elections where voters cast preference ballots, voters often provide only a partial ranking of the candidates. Despite this empirical reality, prior social choice literature frequently analyzes fairness cri…
Fairness