paper-with-me

홈 › Papers

Revisiting Peng's Q($λ$) for Modern Reinforcement Learning

2021-02-27 · Tadashi Kozuno, Yunhao Tang, Mark Rowland, Rémi Munos, Steven Kapturowski, Will Dabney, Michal Valko, David Abel

Off-policy multi-step reinforcement learning algorithms consist of conservative and non-conservative algorithms: the former actively cut traces, whereas the latter do not. Recently, Munos et al. (2016) proved the convergence of conservative algorithms to an optimal Q-function. In contrast, non-conservative algorithms are thought to be unsafe and have a limited or no theoretical guarantee. Nonetheless, recent studies have shown that non-conservative algorithms empirically outperform conservative ones. Motivated by the empirical results and the lack of theory, we carry out theoretical analyses of Peng's Q($\lambda$), a representative example of non-conservative algorithms. We prove that it also converges to an optimal policy provided that the behavior policy slowly tracks a greedy policy in a way similar to conservative policy iteration. Such a result has been conjectured to be true but has not been proven. We also experiment with Peng's Q($\lambda$) in complex continuous control tasks, confirming that Peng's Q($\lambda$) often outperforms conservative algorithms despite its simplicity. These results indicate that Peng's Q($\lambda$), which was thought to be unsafe, is a theoretically-sound and practically effective algorithm.

📄 PDF Abstract BibTeX arXiv:2103.00107

Code (0)

등록된 구현이 없습니다.

Tasks

continuous-controlContinuous Controlreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

OpenGait: Revisiting Gait Recognition Toward Better Practicality

2022-11-12 · Chao Fan, Junhao Liang, Chuanfu Shen, Saihui Hou 외

Gait recognition is one of the most critical long-distance identification technologies and increasingly gains popularity in both research and industry communities. Despite the significant progress made in indoor datasets…

Gait Recognition

OpenGait: Revisiting Gait Recognition Towards Better Practicality

2023-01-01 · CVPR 2023 1 · Chao Fan, Junhao Liang, Chuanfu Shen, Saihui Hou 외

Gait recognition is one of the most critical long-distance identification technologies and increasingly gains popularity in both research and industry communities. Despite the significant progress made in indoor data…

Gait Recognition

Diving with Penguins: Detecting Penguins and their Prey in Animal-borne Underwater Videos via Deep Learning

2023-08-14 · Kejia Zhang, Mingyu Yang, Stephen D. J. Lang, Alistair M. McInnes 외

African penguins (Spheniscus demersus) are an endangered species. Little is known regarding their underwater hunting strategies and associated predation success rates, yet this is essential for guiding conservation. Mode…

ELF OpenGo: An Analysis and Open Reimplementation of AlphaZero

2019-02-12 · Yuandong Tian, Jerry Ma, Qucheng Gong, Shubho Sengupta 외

The AlphaGo, AlphaGo Zero, and AlphaZero series of algorithms are remarkable demonstrations of deep reinforcement learning's capabilities, achieving superhuman performance in the complex game of Go with progressively inc…

Game of Go

PEnGUiN: Partially Equivariant Graph NeUral Networks for Sample Efficient MARL

2025-03-19 · Joshua McClellan, Greyson Brothers, Furong Huang, Pratap Tokekar

Equivariant Graph Neural Networks (EGNNs) have emerged as a promising approach in Multi-Agent Reinforcement Learning (MARL), leveraging symmetry guarantees to greatly improve sample efficiency and generalization. However…

Multi-agent Reinforcement Learning