paper-with-me

홈 › Papers

Avoiding Overfitting to the Importance Weights in Offline Policy Optimization

2021-09-29 · Yao Liu, Emma Brunskill

Offline policy optimization has a critical impact on many real-world decision-making problems, as online learning is costly and concerning in many applications. Importance sampling and its variants are a widely used type of estimator in offline policy evaluation, which can be helpful to remove assumptions on the chosen function approximations used to represent value functions and process models. In this paper, we identify an important overfitting phenomenon in optimizing the importance weighted return, and propose an algorithm to avoid this overfitting. We provide a theoretical justification of the proposed algorithm through a better per-state-neighborhood normalization condition and show the limitation of previous attempts to this approach through an illustrative example. We further test our proposed method in a healthcare-inspired simulator and a logged dataset collected from real hospitals. These experiments show the proposed method with less overfitting and better test performance compared with state-of-the-art batch reinforcement learning algorithms.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Offline Policy Optimization with Eligible Actions

2022-07-01 · Yao Liu, Yannis Flet-Berliac, Emma Brunskill

Offline policy optimization could have a large impact on many real-world decision-making problems, as online learning may be infeasible in many applications. Importance sampling and its variants are a commonly used type …

continuous-controlContinuous ControlDecision Making

Text Generation by Learning from Demonstrations

2020-09-16 · ICLR 2021 1 · Richard Yuanzhe Pang, He He

Current approaches to text generation largely rely on autoregressive models and maximum likelihood estimation. This paradigm leads to (i) diverse but low-quality samples due to mismatched learning objective and evaluatio…

Machine TranslationQuestion GenerationQuestion-GenerationReinforcement Learning (RL)+2

Adaptive Mixture Importance Sampling for Automated Ads Auction Tuning

2024-09-20 · Yimeng Jia, Kaushal Paneri, Rong Huang, Kailash Singh Maurya 외

This paper introduces Adaptive Mixture Importance Sampling (AMIS) as a novel approach for optimizing key performance indicators (KPIs) in large-scale recommender systems, such as online ad auctions. Traditional importanc…

Decision MakingDiversityRecommendation Systems

Balanced off-policy evaluation in general action spaces

2019-06-09 · Arjun Sondhi, David Arbour, Drew Dimmery

Estimation of importance sampling weights for off-policy evaluation of contextual bandits often results in imbalance - a mismatch between the desired and the actual distribution of state-action pairs after weighting. In …

Binary ClassificationcounterfactualMulti-Armed BanditsOff-policy evaluation

Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching

2026-05-07 · Xiang Li, Nan Jiang arxiv

We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one for each data point, such that the reweighted rewards approximate the expected …

Offline RL