Reinforcement learning with restrictions on the action set
Consider a 2-player normal-form game repeated over time. We introduce an adaptive learning procedure, where the players only observe their own realized payoff at each stage. We assume that agents do not know their own payoff function, and have no information on the other player. Furthermore, we assume that they have restrictions on their own action set such that, at each stage, their choice is limited to a subset of their action set. We prove that the empirical distributions of play converge to the set of Nash equilibria for zero-sum and potential games, and games where one player has two actions.
Code (0)
등록된 구현이 없습니다.
Tasks
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Dynamic Interval Restrictions on Action Spaces in Deep Reinforcement Learning for Obstacle Avoidance
Deep reinforcement learning algorithms typically act on the same set of actions. However, this is not sufficient for a wide range of real-world applications where different subsets are available at each step. In this the…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningModeling and Optimization of Epidemiological Control Policies Through Reinforcement Learning
Pandemics involve the high transmission of a disease that impacts global and local health and economic patterns. The impact of a pandemic can be minimized by enforcing certain restrictions on a community. However, while …
Multi-Objective Reinforcement Learningreinforcement-learningReinforcement LearningGeneral policy mapping: online continual reinforcement learning inspired on the insect brain
We have developed a model for online continual or lifelong reinforcement learning (RL) inspired on the insect brain. Our model leverages the offline training of a feature extraction and a common general policy layer to e…
reinforcement-learningReinforcement Learning (RL)Provably Efficient Reinforcement Learning with Aggregated States
We establish that an optimistic variant of Q-learning applied to a fixed-horizon episodic Markov decision process with an aggregated state representation incurs regret $\tilde{\mathcal{O}}(\sqrt{H^5 M K} + \epsilon HK)$,…
Q-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)AirDialogue: An Environment for Goal-Oriented Dialogue Research
Recent progress in dialogue generation has inspired a number of studies on dialogue systems that are capable of accomplishing tasks through natural language interactions. A promising direction among these studies is the …
Dialogue GenerationReinforcement LearningText Generation