paper-with-me

홈 › Papers

Predicting Long Term Sequential Policy Value Using Softer Surrogates

2024-12-30 · HyunJi Nam, Allen Nie, Ge Gao, Vasilis Syrgkanis, Emma Brunskill

Off-policy policy evaluation (OPE) estimates the outcome of a new policy using historical data collected from a different policy. However, existing OPE methods cannot handle cases when the new policy introduces novel actions. This issue commonly occurs in real-world domains, like healthcare, as new drugs and treatments are continuously developed. Novel actions necessitate on-policy data collection, which can be burdensome and expensive if the outcome of interest takes a substantial amount of time to observe--for example, in multi-year clinical trials. This raises a key question of how to predict the long-term outcome of a policy after only observing its short-term effects? Though in general this problem is intractable, under some surrogacy conditions, the short-term on-policy data can be combined with the long-term historical data to make accurate predictions about the new policy's long-term value. In two simulated healthcare examples--HIV and sepsis management--we show that our estimators can provide accurate predictions about the policy value only after observing 10\% of the full horizon data. We also provide finite sample analysis of our doubly robust estimators.

📄 PDF Abstract BibTeX arXiv:2412.20638

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor

2022-06-01 · Wanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng 외

Long-term engagement is preferred over immediate engagement in sequential recommendation as it directly affects product operational metrics such as daily active users (DAUs) and dwell time. Meanwhile, reinforcement learn…

Reinforcement Learning (RL)Sequential Recommendation

Learning-to-defer for sequential medical decision-making under uncertainty

2021-09-13 · Shalmali Joshi, Sonali Parbhoo, Finale Doshi-Velez

Learning-to-defer is a framework to automatically defer decision-making to a human expert when ML-based decisions are deemed unreliable. Existing learning-to-defer frameworks are not designed for sequential settings. Tha…

Decision MakingDecision Making Under UncertaintyModel-based Reinforcement LearningReinforcement Learning (RL)+1

BCRLSP: An Offline Reinforcement Learning Framework for Sequential Targeted Promotion

2022-07-16 · Fanglin Chen, Xiao Liu, Bo Tang, Feiyu Xiong 외

We utilize an offline reinforcement learning (RL) model for sequential targeted promotion in the presence of budget constraints in a real-world business environment. In our application, the mobile app aims to boost custo…

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Adapting Static Fairness to Sequential Decision-Making: Bias Mitigation Strategies towards Equal Long-term Benefit Rate

2023-09-07 · Yuancheng Xu, ChengHao Deng, Yanchao Sun, Ruijie Zheng 외

Decisions made by machine learning models can have lasting impacts, making long-term fairness a critical consideration. It has been observed that ignoring the long-term effect and directly applying fairness criterion in …

Decision MakingFairnessSequential Decision Making

Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback

2024-01-17 · Teng Xiao, Suhang Wang

Probabilistic learning to rank (LTR) has been the dominating approach for optimizing the ranking metric, but cannot maximize long-term rewards. Reinforcement learning models have been proposed to maximize user long-term …

Decision MakingLearning-To-Rankreinforcement-learningSequential Decision Making