paper-with-me

홈 › Papers

SPQR: Controlling Q-ensemble Independence with Spiked Random Model for Reinforcement Learning

2024-01-06 · NeurIPS 2023 11 · Dohyeok Lee, Seungyub Han, Taehyun Cho, Jungwoo Lee

Alleviating overestimation bias is a critical challenge for deep reinforcement learning to achieve successful performance on more complex tasks or offline datasets containing out-of-distribution data. In order to overcome overestimation bias, ensemble methods for Q-learning have been investigated to exploit the diversity of multiple Q-functions. Since network initialization has been the predominant approach to promote diversity in Q-functions, heuristically designed diversity injection methods have been studied in the literature. However, previous studies have not attempted to approach guaranteed independence over an ensemble from a theoretical perspective. By introducing a novel regularization loss for Q-ensemble independence based on random matrix theory, we propose spiked Wishart Q-ensemble independence regularization (SPQR) for reinforcement learning. Specifically, we modify the intractable hypothesis testing criterion for the Q-ensemble independence into a tractable KL divergence between the spectral distribution of the Q-ensemble and the target Wigner's semicircle distribution. We implement SPQR in several online and offline ensemble Q-learning algorithms. In the experiments, SPQR outperforms the baseline algorithms in both online and offline RL benchmarks.

📄 PDF Abstract BibTeX arXiv:2401.03137

Code (1)

dohyeoklee/SPQR 공식 구현 pytorch

Tasks

Deep Reinforcement LearningDiversityOffline RLQ-Learningreinforcement-learningReinforcement Learning

Methods 이 논문이 사용한 방법론

Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…

Similar Papers 제목 키워드 기반

Optimality and Sub-optimality of PCA I: Spiked Random Matrix Models

2018-07-02 · Amelia Perry, Alexander S. Wein, Afonso S. Bandeira, Ankur Moitra

A central problem of random matrix theory is to understand the eigenvalues of spiked random matrix models, introduced by Johnstone, in which a prominent eigenvector (or "spike") is planted into a random matrix. These dis…

Optimality and Sub-optimality of PCA for Spiked Random Matrices and Synchronization

2016-09-19 · Amelia Perry, Alexander S. Wein, Afonso S. Bandeira, Ankur Moitra

A central problem of random matrix theory is to understand the eigenvalues of spiked random matrix models, in which a prominent eigenvector is planted into a random matrix. These distributions form natural statistical mo…

The Binary Expansion Randomized Ensemble Test (BERET)

2019-12-08 · Duyeol Lee, Kai Zhang, Michael R. Kosorok

Recently, the binary expansion testing framework was introduced to test the independence of two continuous random variables by utilizing symmetry statistics that are complete sufficient statistics for dependence. We deve…

Machine Learning for RealisticBall Detection in RoboCup SPL

2017-07-12 · Domenico Bloisi, Francesco Del Duchetto, Tiziano Manoni, Vincenzo Suriani

In this technical report, we describe the use of a machine learning approach for detecting the realistic black and white ball currently in use in the RoboCup Standard Platform League. Our aim is to provide a ready-to-use…

BIG-bench Machine Learning

SPQR: An R Package for Semi-Parametric Density and Quantile Regression

2022-10-26 · Steven G. Xu, Reetam Majumder, Brian J. Reich

We develop an R package SPQR that implements the semi-parametric quantile regression (SPQR) method in Xu and Reich (2021). The method begins by fitting a flexible density regression model using monotonic splines whose we…

quantile regressionregression