paper-with-me

홈 › Papers

Log-normality and Skewness of Estimated State/Action Values in Reinforcement Learning

2017-12-01 · NeurIPS 2017 12 · Liangpeng Zhang, Ke Tang, Xin Yao

Under/overestimation of state/action values are harmful for reinforcement learning agents. In this paper, we show that a state/action value estimated using the Bellman equation can be decomposed to a weighted sum of path-wise values that follow log-normal distributions. Since log-normal distributions are skewed, the distribution of estimated state/action values can also be skewed, leading to an imbalanced likelihood of under/overestimation. The degree of such imbalance can vary greatly among actions and policies within a single problem instance, making the agent prone to select actions/policies that have inferior expected return and higher likelihood of overestimation. We present a comprehensive analysis to such skewness, examine its factors and impacts through both theoretical and empirical results, and discuss the possible ways to reduce its undesirable effects.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

NPSA: Nonorthogonal Principal Skewness Analysis

2019-07-23 · Xiurui Geng, Lei Wang

Principal skewness analysis (PSA) has been introduced for feature extraction in hyperspectral imagery. As a third-order generalization of principal component analysis (PCA), its solution of searching for the locally maxi…

Variation-Incentive Loss Re-weighting for Regression Analysis on Biased Data

2021-09-14 · Wentai Wu, Ligang He, Weiwei Lin

Both classification and regression tasks are susceptible to the biased distribution of training data. However, existing approaches are focused on the class-imbalanced learning and cannot be applied to the problems of num…

regression

Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning

2024-03-12 · Motoki Omura, Takayuki Osa, Yusuke Mukuta, Tatsuya Harada

In deep reinforcement learning, estimating the value function to evaluate the quality of states and actions is essential. The value function is often trained using the least squares method, which implicitly assumes a Gau…

continuous-controlContinuous ControlDeep Reinforcement LearningMuJoCo+3

Monotonicity-Constrained Nonparametric Estimation and Inference for First-Price Auctions

2019-09-27

We propose a new nonparametric estimator for first-price auctions with independent private values that imposes the monotonicity constraint on the estimated inverse bidding strategy. We show that our estimator has a small…

The impact of COVID-19 on the stock market crash risk in China

2020-09-17 · Zhifeng Liu, Toan Luu Duc Huynh, Peng-Fei Dai

This study investigates the impact of the COVID-19 pandemic on the stock market crash risk in China. For this purpose, we first estimated the conditional skewness of the return distribution from a GARCH with skewness (GA…