paper-with-me

홈 › Papers

On the Equilibrium between Feasible Zone and Uncertain Model in Safe Exploration

2026-01-31 · Yujie Yang, Zhilong Zheng, Shengbo Eben Li arxiv

Ensuring the safety of environmental exploration is a critical problem in reinforcement learning (RL). While limiting exploration to a feasible zone has become widely accepted as a way to ensure safety, key questions remain unresolved: what is the maximum feasible zone achievable through exploration, and how can it be identified? This paper, for the first time, answers these questions by revealing that the goal of safe exploration is to find the equilibrium between the feasible zone and the environment model. This conclusion is based on the understanding that these two components are interdependent: a larger feasible zone leads to a more accurate environment model, and a more accurate model, in turn, enables exploring a larger zone. We propose the first equilibrium-oriented safe exploration framework called safe equilibrium exploration (SEE), which alternates between finding the maximum feasible zone and the least uncertain model. Using a graph formulation of the uncertain model, we prove that the uncertain model obtained by SEE is monotonically refined, the feasible zones monotonically expand, and both converge to the equilibrium of safe exploration. Experiments on classic control tasks show that our algorithm successfully expands the feasible zones with zero constraint violation, and achieves the equilibrium of safe exploration within a few iterations.

📄 PDF Abstract BibTeX arXiv:2602.00636

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Uncertainty Quantification for Surface Ozone Emulators using Deep Learning

2025-08-06 · Kelsey Doerksen, Yuliya Marchetti, Steven Lu, Kevin Bowman 외 arxiv

Air pollution is a global hazard, and as of 2023, 94\% of the world's population is exposed to unsafe pollution levels. Surface Ozone (O3), an important pollutant, and the drivers of its trends are difficult to model, an…

Decision Making

A Comparison of Reinforcement Learning and Optimal Control Methods for Path Planning

2026-04-14 · Qiang Le, Yaguang Yang, Isaac E. Weintraub arxiv

Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge. While traditional optimal control methods can find ideal paths, the computational time is often too slow for real-time decisi…

Reinforcement LearningAutonomous Vehicles

A Chance-Constrained Stochastic Electricity Market

2019-12-18

Efficiently accommodating uncertain renewable resources in wholesale electricity markets is among the foremost priorities of market regulators in the US, UK and EU nations. However, existing deterministic market designs …

Probabilistic Weapon Engagement Zones for a Turn Constrained Pursuer

2025-12-05 · Grant Stagg, Isaac E. Weintraub, Cameron K. Peterson arxiv

Curve-straight probabilistic engagement zones (CSPEZ) quantify the spatial regions an evader should avoid to reduce capture risk from a turn-rate-limited pursuer following a curve-straight path with uncertain parameters …

ROSA-RL: Uncertainty-Aware Roundabout Optimized Speed Advisory with Reinforcement Learning

2026-06-15 · Anna-Lena Schlamp, Jeremias Gerner, Klaus Bogenberger, Werner Huber 외 arxiv

Roundabouts challenge automated driving in mixed traffic, as heterogeneous and non-deterministic human behavior, unknown driving intentions, and high interaction complexity create uncertainty about whether the conflict z…

Reinforcement Learning