paper-with-me

Papers

Provable Multi-Objective Reinforcement Learning with Generative Models

2020-11-19 · Dongruo Zhou, Jiahao Chen, Quanquan Gu

Multi-objective reinforcement learning (MORL) is an extension of ordinary, single-objective reinforcement learning (RL) that is applicable to many real-world tasks where multiple objectives exist without known relative costs. We study the problem of single policy MORL, which learns an optimal policy given the preference of objectives. Existing methods require strong assumptions such as exact knowledge of the multi-objective Markov decision process, and are analyzed in the limit of infinite data and time. We propose a new algorithm called model-based envelop value iteration (EVI), which generalizes the enveloped multi-objective $Q$-learning algorithm in Yang et al., 2019. Our method can learn a near-optimal value function with polynomial sample complexity and linear convergence speed. To the best of our knowledge, this is the first finite-sample analysis of MORL algorithms.

📄 PDF Abstract BibTeX arXiv:2011.10134

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Objective Reinforcement LearningQ-Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Multi-objective Reinforcement Learning with Nonlinear Preferences: Provable Approximation for Maximizing Expected Scalarized Return

2023-11-05 · Nianli Peng, Muhang Tian, Brandon Fain

We study multi-objective reinforcement learning with nonlinear preferences over trajectories. That is, we maximize the expected value of a nonlinear function over accumulated rewards (expected scalarized return or ESR) i…

FairnessMulti-Objective Reinforcement Learningreinforcement-learning

Conditional Generative Quantile Networks via Optimal Transport and Convex Potentials

2021-09-29 · Jesse Sun, Dihong Jiang, YaoLiang Yu

Quantile regression has a natural extension to generative modelling by leveraging a stronger convergence in pointwise rather than in distribution. While the pinball quantile loss works in the scalar case, it does not hav…

quantile regression

Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning

2024-05-05 · Tianchen Zhou, FNU Hairi, Haibo Yang, Jia Liu 외

Reinforcement learning with multiple, potentially conflicting objectives is pervasive in real-world applications, while this problem remains theoretically under-explored. This paper tackles the multi-objective reinforcem…

Multi-Objective Reinforcement Learningreinforcement-learningReinforcement Learning

Adversarial Training and Provable Robustness: A Tale of Two Objectives

2020-08-13 · Jiameng Fan, Wenchao Li

We propose a principled framework that combines adversarial training and provable robustness verification for training certifiably robust neural networks. We formulate the training problem as a joint optimization problem…

Vocal Bursts Valence Prediction

Causal-Aware Generative Adversarial Networks with Reinforcement Learning

2025-10-28 · Tu Anh Hoang Nguyen, Dang Nguyen, Tri-Nhan Vo, Thuc Duy Le 외 arxiv

The utility of tabular data for tasks ranging from model training to large-scale data analysis is often constrained by privacy concerns or regulatory hurdles. While existing data generation methods, particularly those ba…

Reinforcement Learning