paper-with-me

홈 › Papers

Minimax Weight and Q-Function Learning for Off-Policy Evaluation

2019-10-28 · ICML 2020 1 · Masatoshi Uehara, Jiawei Huang, Nan Jiang

We provide theoretical investigations into off-policy evaluation in reinforcement learning using function approximators for (marginalized) importance weights and value functions. Our contributions include: (1) A new estimator, MWL, that directly estimates importance ratios over the state-action distributions, removing the reliance on knowledge of the behavior policy as in prior work (Liu et al., 2018). (2) Another new estimator, MQL, obtained by swapping the roles of importance weights and value-functions in MWL. MQL has an intuitive interpretation of minimizing average Bellman errors and can be combined with MWL in a doubly robust manner. (3) Several additional results that offer further insights into these methods, including the sample complexity analyses of MWL and MQL, their asymptotic optimality in the tabular setting, how the learned importance weights depend the choice of the discriminator class, and how our methods provide a unified view of some old and new algorithms in RL.

📄 PDF Abstract BibTeX arXiv:1910.12809

Code (0)

등록된 구현이 없습니다.

Tasks

Off-policy evaluationReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Minimax Value Interval for Off-Policy Evaluation and Policy Optimization

2020-02-06 · NeurIPS 2020 12 · Nan Jiang, Jiawei Huang

We study minimax methods for off-policy evaluation (OPE) using value functions and marginalized importance weights. Despite that they hold promises of overcoming the exponential variance in traditional importance samplin…

Efficient ExplorationOff-policy evaluationvalid

Finite Sample Analysis of Minimax Offline Reinforcement Learning: Completeness, Fast Rates and First-Order Efficiency

2021-02-05 · Masatoshi Uehara, Masaaki Imaizumi, Nan Jiang, Nathan Kallus 외

We offer a theoretical characterization of off-policy evaluation (OPE) in reinforcement learning using function approximation for marginal importance weights and $q$-functions when these are estimated using recent minima…

Off-policy evaluationreinforcement-learningReinforcement Learning (RL)

Minimax-Optimal Off-Policy Evaluation with Linear Function Approximation

2020-02-21 · ICML 2020 1 · Yaqi Duan, Mengdi Wang

This paper studies the statistical theory of batch data reinforcement learning with function approximation. Consider the off-policy evaluation problem, which is to estimate the cumulative value of a new target policy fro…

Off-policy evaluationReinforcement Learning

A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision Processes

2021-11-12 · Chengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan Jiang

We consider off-policy evaluation (OPE) in Partially Observable Markov Decision Processes (POMDPs), where the evaluation policy depends only on observable variables and the behavior policy depends on unobservable latent …

Off-policy evaluation

Minimax Least-Square Policy Iteration for Cost-Aware Defense of Traffic Routing against Unknown Threats

2024-04-07 · Yuzhen Zhan, Li Jin

Dynamic routing is one of the representative control scheme in transportation, production lines, and data transmission. In the modern context of connectivity and autonomy, routing decisions are potentially vulnerable to …