Provable Multi-Objective Reinforcement Learning with Generative Models
Multi-objective reinforcement learning (MORL) is an extension of ordinary, single-objective reinforcement learning (RL) that is applicable to many real-world tasks where multiple objectives exist without known relative costs. We study the problem of single policy MORL, which learns an optimal policy given the preference of objectives. Existing methods require strong assumptions such as exact knowledge of the multi-objective Markov decision process, and are analyzed in the limit of infinite data and time. We propose a new algorithm called model-based envelop value iteration (EVI), which generalizes the enveloped multi-objective $Q$-learning algorithm in Yang et al., 2019. Our method can learn a near-optimal value function with polynomial sample complexity and linear convergence speed. To the best of our knowledge, this is the first finite-sample analysis of MORL algorithms.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-Objective Reinforcement LearningQ-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Multi-objective Reinforcement Learning with Nonlinear Preferences: Provable Approximation for Maximizing Expected Scalarized Return
We study multi-objective reinforcement learning with nonlinear preferences over trajectories. That is, we maximize the expected value of a nonlinear function over accumulated rewards (expected scalarized return or ESR) i…
FairnessMulti-Objective Reinforcement Learningreinforcement-learningConditional Generative Quantile Networks via Optimal Transport and Convex Potentials
Quantile regression has a natural extension to generative modelling by leveraging a stronger convergence in pointwise rather than in distribution. While the pinball quantile loss works in the scalar case, it does not hav…
quantile regressionFinite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning
Reinforcement learning with multiple, potentially conflicting objectives is pervasive in real-world applications, while this problem remains theoretically under-explored. This paper tackles the multi-objective reinforcem…
Multi-Objective Reinforcement Learningreinforcement-learningReinforcement LearningAdversarial Training and Provable Robustness: A Tale of Two Objectives
We propose a principled framework that combines adversarial training and provable robustness verification for training certifiably robust neural networks. We formulate the training problem as a joint optimization problem…
Vocal Bursts Valence PredictionCausal-Aware Generative Adversarial Networks with Reinforcement Learning
The utility of tabular data for tasks ranging from model training to large-scale data analysis is often constrained by privacy concerns or regulatory hurdles. While existing data generation methods, particularly those ba…
Reinforcement Learning