paper-with-me

홈 › Papers

The State-Action-Reward-State-Action Algorithm in Spatial Prisoner's Dilemma Game

2024-06-25 · Lanyu Yang, Dongchun Jiang, Fuqiang Guo, Mingjian Fu

Cooperative behavior is prevalent in both human society and nature. Understanding the emergence and maintenance of cooperation among self-interested individuals remains a significant challenge in evolutionary biology and social sciences. Reinforcement learning (RL) provides a suitable framework for studying evolutionary game theory as it can adapt to environmental changes and maximize expected benefits. In this study, we employ the State-Action-Reward-State-Action (SARSA) algorithm as the decision-making mechanism for individuals in evolutionary game theory. Initially, we apply SARSA to imitation learning, where agents select neighbors to imitate based on rewards. This approach allows us to observe behavioral changes in agents without independent decision-making abilities. Subsequently, SARSA is utilized for primary agents to independently choose cooperation or betrayal with their neighbors. We evaluate the impact of SARSA on cooperation rates by analyzing variations in rewards and the distribution of cooperators and defectors within the network.

📄 PDF Abstract BibTeX arXiv:2406.17326

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingImitation LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Sarsa Sarsa is an on-policy TD control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} + \gamma{Q}\left(S\_{t+1},…

Similar Papers 제목 키워드 기반

Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures

2024-12-09 · Adrien Bolland, Gaspard Lambrechts, Damien Ernst

We introduce a new maximum entropy reinforcement learning framework based on the distribution of states and actions visited by a policy. More precisely, an intrinsic reward function is added to the reward function of the…

reinforcement-learningReinforcement Learning

Learning Symbolic Representations for Reinforcement Learning of Non-Markovian Behavior

2023-01-08 · Phillip J. K. Christoffersen, Andrew C. Li, Rodrigo Toro Icarte, Sheila A. McIlraith

Many real-world reinforcement learning (RL) problems necessitate learning complex, temporally extended behavior that may only receive reward signal when the behavior is completed. If the reward-worthy behavior is known, …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Transition-based versus State-based Reward Functions for MDPs with Value-at-Risk

2016-12-07 · Shuai Ma, Jia Yuan Yu

In reinforcement learning, the reward function on current state and action is widely used. When the objective is about the expectation of the (discounted) total reward only, it works perfectly. However, if the objective …

Reinforcement Learning

Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics

2025-10-17 · Sunmook Choi, Yahya Sattar, Yassir Jedra, Maryam Fazel 외 arxiv

We study a nonstationary bandit problem where rewards depend on both actions and latent states, the latter governed by unknown linear dynamics. Crucially, the state dynamics also depend on the actions, resulting in tensi…

AUPO -- Abstracted Until Proven Otherwise: A Reward Distribution Based Abstraction Algorithm

2025-10-27 · Robin Schmöcker, Alexander Dockhorn, Bodo Rosenhahn arxiv

We introduce a novel, drop-in modification to Monte Carlo Tree Search's (MCTS) decision policy that we call AUPO. Comparisons based on a range of IPPC benchmark problems show that AUPO clearly outperforms MCTS. AUPO is a…