paper-with-me

홈 › Papers

Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations

2025-12-12 · S. R. Eshwar, Aniruddha Mukherjee, Kintan Saha, Krishna Agarwal, Gugan Thoppe, Aditya Gopalan, Gal Dalal arxiv

In a recent work, we proposed Reliable Policy Iteration (RPI), that restores policy iteration's monotonicity-of-value-estimates property to the function approximation setting. Here, we assess the robustness of RPI's empirical performance on two classical control tasks -- CartPole and Inverted Pendulum -- under changes to neural network and environmental parameters. Relative to DQN, Double DQN, DDPG, TD3, and PPO, RPI reaches near-optimal performance early and sustains this policy as training proceeds. Because deep RL methods are often hampered by sample inefficiency, training instability, and hyperparameter sensitivity, our results highlight RPI's promise as a more reliable alternative.

📄 PDF Abstract BibTeX arXiv:2512.12088

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Convergence of Modified Policy Iteration in Risk Sensitive Exponential Cost Markov Decision Processes

2023-02-08 · Yashaswini Murthy, Mehrdad Moharrami, R. Srikant

Modified policy iteration (MPI) is a dynamic programming algorithm that combines elements of policy iteration and value iteration. The convergence of MPI has been well studied in the context of discounted and average-cos…

Computational Efficiency

Robust Regularized Policy Iteration under Transition Uncertainty

2026-03-10 · Hongqiang Lin, Zhenghui Fu, Weihao Tang, Pengfei Wang 외 arxiv

Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift. The learned policy may visit out-of-distributio…

Reinforcement LearningOffline RL

Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization

2025-07-04 · Buqing Nie, Yangqing Fu, Jingtian Ji, Yue Gao arxiv

Reinforcement Learning (RL) has achieved remarkable success in sequential decision tasks. However, recent studies have revealed the vulnerability of RL policies to different perturbations, raising concerns about their ef…

Reinforcement Learning

Towards Efficient Risk-Sensitive Policy Gradient: An Iteration Complexity Analysis

2024-03-13 · Rui Liu, Anish Gupta, Erfaun Noorani, Pratap Tokekar

Reinforcement Learning (RL) has shown exceptional performance across various applications, enabling autonomous agents to learn optimal policies through interaction with their environments. However, traditional RL framewo…

Policy Gradient MethodsReinforcement Learning (RL)Robot Navigation

Learning Robust Options

2018-02-09 · Daniel J. Mankowitz, Timothy A. Mann, Pierre-Luc Bacon, Doina Precup 외

Robust reinforcement learning aims to produce policies that have strong guarantees even in the face of environments/transition models whose parameters have strong uncertainty. Existing work uses value-based methods and t…

Reinforcement Learning