paper-with-me

Papers

Improving Generalization in Mountain Car Through the Partitioned Parameterized Policy Approach via Quasi-Stochastic Gradient Descent

2021-05-28 · Caleb M. Bowyer

The reinforcement learning problem of finding a control policy that minimizes the minimum time objective for the Mountain Car environment is considered. Particularly, a class of parameterized nonlinear feedback policies is optimized over to reach the top of the highest mountain peak in minimum time. The optimization is carried out using quasi-Stochastic Gradient Descent (qSGD) methods. In attempting to find the optimal minimum time policy, a new parameterized policy approach is considered that seeks to learn an optimal policy parameter for different regions of the state space, rather than rely on a single macroscopic policy parameter for the entire state space. This partitioned parameterized policy approach is shown to outperform the uniform parameterized policy approach and lead to greater generalization than prior methods, where the Mountain Car became trapped in circular trajectories in the state space.

📄 PDF Abstract BibTeX arXiv:2105.13986

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deterministic Policy Optimization by Combining Pathwise and Score Function Estimators for Discrete Action Spaces

2017-11-21 · Daniel Levy, Stefano Ermon

Policy optimization methods have shown great promise in solving complex reinforcement and imitation learning tasks. While model-free methods are broadly applicable, they often require many samples to optimize complex pol…

AcrobotImitation Learning

Representation Learning for Continuous Action Spaces is Beneficial for Efficient Policy Learning

2022-11-23 · Tingting Zhao, Ying Wang, Wei Sun, Yarui Chen 외

Deep reinforcement learning (DRL) breaks through the bottlenecks of traditional reinforcement learning (RL) with the help of the perception capability of deep learning and has been widely applied in real-world problems.W…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

Proximal Policy Optimization with Evolutionary Mutations

2026-01-21 · Casimir Czworkowski, Stephen Hornish, Alhassan S. Yasin arxiv

Proximal Policy Optimization (PPO) is a widely used reinforcement learning algorithm known for its stability and sample efficiency, but it often suffers from premature convergence due to limited exploration. In this pape…

Reinforcement LearningOpenAI Gym

Biological barriers to forest pest invasions: A novel host tree slows mountain pine beetle range expansion

2024-12-11 · Evan C. Johnson, Antonia Musso, Catherine Cullingham, Mark A. Lewis

Following widespread outbreaks across western North America, mountain pine beetle recently expanded its range from British Columbia into Alberta. However, mountain pine beetle's eastward expansion across Canada has stall…

Snowpack Estimation in Key Mountainous Water Basins from Openly-Available, Multimodal Data Sources

2022-08-08 · Malachy Moran, Kayla Woputz, Derrick Hee, Manuela Girotto 외

Accurately estimating the snowpack in key mountainous basins is critical for water resource managers to make decisions that impact local and global economies, wildlife, and public policy. Currently, this estimation requi…