paper-with-me

홈 › Papers

On the Optimality of Batch Policy Optimization Algorithms

2021-04-06 · Chenjun Xiao, Yifan Wu, Tor Lattimore, Bo Dai, Jincheng Mei, Lihong Li, Csaba Szepesvari, Dale Schuurmans

Batch policy optimization considers leveraging existing data for policy construction before interacting with an environment. Although interest in this problem has grown significantly in recent years, its theoretical foundations remain under-developed. To advance the understanding of this problem, we provide three results that characterize the limits and possibilities of batch policy optimization in the finite-armed stochastic bandit setting. First, we introduce a class of confidence-adjusted index algorithms that unifies optimistic and pessimistic principles in a common framework, which enables a general analysis. For this family, we show that any confidence-adjusted index algorithm is minimax optimal, whether it be optimistic, pessimistic or neutral. Our analysis reveals that instance-dependent optimality, commonly used to establish optimality of on-line stochastic bandit algorithms, cannot be achieved by any algorithm in the batch setting. In particular, for any algorithm that performs optimally in some environment, there exists another environment where the same algorithm suffers arbitrarily larger regret. Therefore, to establish a framework for distinguishing algorithms, we introduce a new weighted-minimax criterion that considers the inherent difficulty of optimal value prediction. We demonstrate how this criterion can be used to justify commonly used pessimistic principles for batch policy optimization.

📄 PDF Abstract BibTeX arXiv:2104.02293

Code (0)

등록된 구현이 없습니다.

Tasks

Value prediction

Similar Papers 제목 키워드 기반

Provably Good Batch Reinforcement Learning Without Great Exploration

2020-07-16 · Yao Liu, Adith Swaminathan, Alekh Agarwal, Emma Brunskill

Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is challenging: a new decision policy may visit …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Provably Good Batch Off-Policy Reinforcement Learning Without Great Exploration

2020-12-01 · NeurIPS 2020 12 · Yao Liu, Adith Swaminathan, Alekh Agarwal, Emma Brunskill

Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is challenging: a new decision policy may visit …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Batch size-invariance for policy optimization

2021-10-01 · Jacob Hilton, Karl Cobbe, John Schulman

We say an algorithm is batch size-invariant if changes to the batch size can largely be compensated for by changes to other hyperparameters. Stochastic gradient descent is well-known to have this property at small batch …

Policy Poisoning in Batch Reinforcement Learning and Control

2019-10-13 · NeurIPS 2019 12 · Yuzhe Ma, Xuezhou Zhang, Wen Sun, Xiaojin Zhu

We study a security threat to batch reinforcement learning and control where the attacker aims to poison the learned policy. The victim is a reinforcement learner / controller which first estimates the dynamics and the r…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Model Selection in Batch Policy Optimization

2021-12-23 · Jonathan N. Lee, George Tucker, Ofir Nachum, Bo Dai

We study the problem of model selection in batch policy optimization: given a fixed, partial-feedback dataset and $M$ model classes, learn a policy with performance that is competitive with the policy derived from the be…

modelModel Selection