paper-with-me

Papers

Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model

2024-02-12 · Mark Rowland, Li Kevin Wenliang, Rémi Munos, Clare Lyle, Yunhao Tang, Will Dabney

We propose a new algorithm for model-based distributional reinforcement learning (RL), and prove that it is minimax-optimal for approximating return distributions with a generative model (up to logarithmic factors), resolving an open question of Zhang et al. (2023). Our analysis provides new theoretical results on categorical approaches to distributional RL, and also introduces a new distributional Bellman equation, the stochastic categorical CDF Bellman equation, which we expect to be of independent interest. We also provide an experimental study comparing several model-based distributional RL algorithms, with several takeaways for practitioners.

📄 PDF Abstract BibTeX arXiv:2402.07598

Code (0)

등록된 구현이 없습니다.

Tasks

Distributional Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Improving Generalization of Reinforcement Learning with Minimax Distributional Soft Actor-Critic

2020-02-13 · Yangang Ren, Jingliang Duan, Shengbo Eben Li, Yang Guan 외

Reinforcement learning (RL) has achieved remarkable performance in numerous sequential decision making and control tasks. However, a common problem is that learned nearly optimal policy always overfits to the training en…

Autonomous DrivingAutonomous VehiclesDecision Makingreinforcement-learning+3

ORVIT: Near-Optimal Online Distributionally Robust Reinforcement Learning

2025-08-05 · Debamita Ghosh, George K. Atia, Yue Wang arxiv

We investigate reinforcement learning (RL) in the presence of distributional mismatch between training and deployment, where policies trained in simulators often underperform in practice due to mismatches between trainin…

Reinforcement Learning

KL-Entropy-Regularized RL with a Generative Model is Minimax Optimal

2022-05-27 · Tadashi Kozuno, Wenhao Yang, Nino Vieillard, Toshinori Kitamura 외

In this work, we consider and analyze the sample complexity of model-free reinforcement learning with a generative model. Particularly, we analyze mirror descent value iteration (MDVI) by Geist et al. (2019) and Vieillar…

reinforcement-learningReinforcement Learning (RL)

Minimax Optimal and Computationally Efficient Algorithms for Distributionally Robust Offline Reinforcement Learning

2024-03-14 · Zhishuai Liu, Pan Xu

Distributionally robust offline reinforcement learning (RL), which seeks robust policy training against environment perturbation by modeling dynamics uncertainty, calls for function approximations when facing large state…

Offline RLReinforcement Learning (RL)

The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model

2023-05-26 · NeurIPS 2023 11 · Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen 외

This paper investigates model robustness in reinforcement learning (RL) to reduce the sim-to-real gap in practice. We adopt the framework of distributionally robust Markov decision processes (RMDPs), aimed at learning a …

Reinforcement Learning (RL)