Model-Based Reinforcement Learning via Stochastic Hybrid Models
Optimal control of general nonlinear systems is a central challenge in automation. Enabled by powerful function approximators, data-driven approaches to control have recently successfully tackled challenging applications. However, such methods often obscure the structure of dynamics and control behind black-box over-parameterized representations, thus limiting our ability to understand closed-loop behavior. This paper adopts a hybrid-system view of nonlinear modeling and control that lends an explicit hierarchical structure to the problem and breaks down complex dynamics into simpler localized units. We consider a sequence modeling paradigm that captures the temporal structure of the data and derive an expectation-maximization (EM) algorithm that automatically decomposes nonlinear dynamics into stochastic piecewise affine models with nonlinear transition boundaries. Furthermore, we show that these time-series models naturally admit a closed-loop extension that we use to extract local polynomial feedback controllers from nonlinear experts via behavioral cloning. Finally, we introduce a novel hybrid relative entropy policy search (Hb-REPS) technique that incorporates the hierarchical nature of hybrid models and optimizes a set of time-invariant piecewise feedback controllers derived from a piecewise polynomial approximation of a global state-value function.
Code (0)
등록된 구현이 없습니다.
Tasks
Imitation LearningmodelModel-based Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Time SeriesTime Series AnalysisSimilar Papers 제목 키워드 기반
Optimising Stochastic Routing for Taxi Fleets with Model Enhanced Reinforcement Learning
The future of mobility-as-a-Service (Maas)should embrace an integrated system of ride-hailing, street-hailing and ride-sharing with optimised intelligent vehicle routing in response to a real-time, stochastic demand patt…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)A Hybrid Stochastic Policy Gradient Algorithm for Reinforcement Learning
We propose a novel hybrid stochastic policy gradient estimator by combining an unbiased policy gradient estimator, the REINFORCE estimator, with another biased one, an adapted SARAH estimator for policy optimization. The…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Opportunities of Hybrid Model-based Reinforcement Learning for Cell Therapy Manufacturing Process Control
Driven by the key challenges of cell therapy manufacturing, including high complexity, high uncertainty, and very limited process observations, we propose a hybrid model-based reinforcement learning (RL) to efficiently g…
Decision MakingModel-based Reinforcement LearningReinforcement Learning (RL)Stochastic OptimizationStochastic Constraint Programming as Reinforcement Learning
Stochastic Constraint Programming (SCP) is an extension of Constraint Programming (CP) used for modelling and solving problems involving constraints and uncertainty. SCP inherits excellent modelling abilities and filteri…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Model-Based Episodic Memory Induces Dynamic Hybrid Controls
Episodic control enables sample efficiency in reinforcement learning by recalling past experiences from an episodic memory. We propose a new model-based episodic memory of trajectories addressing current limitations of e…
modelreinforcement-learningReinforcement LearningReinforcement Learning (RL)