paper-with-me

홈 › Papers

Deep SPI: Safe Policy Improvement via World Models

2025-10-14 · Florent Delgrange, Raphael Avalos, Willem Röpke arxiv

Safe policy improvement (SPI) offers theoretical control over policy updates, yet existing guarantees largely concern offline, tabular reinforcement learning (RL). We study SPI in general online settings, when combined with world model and representation learning. We develop a theoretical framework showing that restricting policy updates to a well-defined neighborhood of the current policy ensures monotonic improvement and convergence. This analysis links transition and reward prediction losses to representation quality, yielding online, "deep" analogues of classical SPI theorems from the offline RL literature. Building on these results, we introduce DeepSPI, a principled on-policy algorithm that couples local transition and reward losses with regularised policy updates. On the ALE-57 benchmark, DeepSPI matches or exceeds strong baselines, including PPO and DeepMDPs, while retaining theoretical guarantees.

📄 PDF Abstract BibTeX arXiv:2510.12312

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningReinforcement LearningOffline RL

Similar Papers 제목 키워드 기반

Safe Policy Learning from Observations

2018-09-27 · Elad Sarafian, Aviv Tamar, Sarit Kraus

In this paper, we consider the problem of learning a policy by observing numerous non-expert agents. Our goal is to extract a policy that, with high-confidence, acts better than the agents' average performance. Such a se…

Safe Policy Improvement with an Estimated Baseline Policy

2019-09-11 · Thiago D. Simão, Romain Laroche, Rémi Tachet des Combes

Previous work has shown the unreliability of existing algorithms in the batch Reinforcement Learning setting, and proposed the theoretically-grounded Safe Policy Improvement with Baseline Bootstrapping (SPIBB) fix: repro…

ManagementReinforcement Learning

Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs

2021-05-31 · NeurIPS 2021 12 · Harsh Satija, Philip S. Thomas, Joelle Pineau, Romain Laroche

We study the problem of Safe Policy Improvement (SPI) under constraints in the offline Reinforcement Learning (RL) setting. We consider the scenario where: (i) we have a dataset collected under a known baseline policy, (…

Reinforcement Learning (RL)

Feasible Policy Iteration for Safe Reinforcement Learning

2023-04-18 · Yujie Yang, Zhilong Zheng, Shengbo Eben Li, Wei Xu 외

Safety is the priority concern when applying reinforcement learning (RL) algorithms to real-world control problems. While policy iteration provides a fundamental algorithm for standard RL, an analogous theoretical algori…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement Learning

Safe Planning and Policy Optimization via World Model Learning

2025-06-05 · Artem Latyshev, Gregory Gorbov, Aleksandr I. Panov

Reinforcement Learning (RL) applications in real-world scenarios must prioritize safety and reliability, which impose strict constraints on agent behavior. Model-based RL leverages predictive world models for action plan…

continuous-controlContinuous ControlReinforcement Learning (RL)