Policy Learning for Balancing Short-Term and Long-Term Rewards
Empirical researchers and decision-makers spanning various domains frequently seek profound insights into the long-term impacts of interventions. While the significance of long-term outcomes is undeniable, an overemphasis on them may inadvertently overshadow short-term gains. Motivated by this, this paper formalizes a new framework for learning the optimal policy that effectively balances both long-term and short-term rewards, where some long-term outcomes are allowed to be missing. In particular, we first present the identifiability of both rewards under mild assumptions. Next, we deduce the semiparametric efficiency bounds, along with the consistency and asymptotic normality of their estimators. We also reveal that short-term outcomes, if associated, contribute to improving the estimator of the long-term reward. Based on the proposed estimators, we develop a principled policy learning approach and further derive the convergence rates of regret and estimation errors associated with the learned policy. Extensive experiments are conducted to validate the effectiveness of the proposed method, demonstrating its practical applicability.
Code (1)
Similar Papers 제목 키워드 기반
Pareto-Optimal Estimation and Policy Learning on Short-term and Long-term Treatment Effects
This paper focuses on developing Pareto-optimal estimation and policy learning to identify the most effective treatment that maximizes the total reward from both short-term and long-term effects, which might conflict wit…
Representation LearningRepresentation Balancing Offline Model-based Reinforcement Learning
One of the main challenges in offline and off-policy reinforcement learning is to cope with the distribution shift that arises from the mismatch between the target policy and the data collection policy. In this paper, we…
modelModel-based Reinforcement LearningOffline RLreinforcement-learning+2Unveiling User Satisfaction and Creator Productivity Trade-Offs in Recommendation Platforms
On User-Generated Content (UGC) platforms, recommendation algorithms significantly impact creators' motivation to produce content as they compete for algorithmically allocated user traffic. This phenomenon subtly shapes …
DiversityLongShortNet: Exploring Temporal and Semantic Features Fusion in Streaming Perception
Streaming perception is a critical task in autonomous driving that requires balancing the latency and accuracy of the autopilot system. However, current methods for streaming perception are limited as they only rely on t…
Autonomous DrivingSTOPS: Short-Term-based Volatility-controlled Policy Search and its Global Convergence
It remains challenging to deploy existing risk-averse approaches to real-world applications. The reasons are multi-fold, including the lack of global optimality guarantee and the necessity of learning from long-term cons…
MuJoCo