Simulator-Driven Deceptive Control via Path Integral Approach
We consider a setting where a supervisor delegates an agent to perform a certain control task, while the agent is incentivized to deviate from the given policy to achieve its own goal. In this work, we synthesize the optimal deceptive policies for an agent who attempts to hide its deviations from the supervisor's policy. We study the deception problem in the continuous-state discrete-time stochastic dynamics setting and, using motivations from hypothesis testing theory, formulate a Kullback-Leibler control problem for the synthesis of deceptive policies. This problem can be solved using backward dynamic programming in principle, which suffers from the curse of dimensionality. However, under the assumption of deterministic state dynamics, we show that the optimal deceptive actions can be generated using path integral control. This allows the agent to numerically compute the deceptive actions via Monte Carlo simulations. Since Monte Carlo simulations can be efficiently parallelized, our approach allows the agent to generate deceptive control actions online. We show that the proposed simulation-driven control approach asymptotically converges to the optimal control distribution.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
CableRobotGraphSim: A Graph Neural Network for Modeling Partially Observable Cable-Driven Robot Dynamics
General-purpose simulators have accelerated the development of robots. Traditional simulators based on first-principles, however, typically require full-state observability or depend on parameter search for system identi…
Graph Neural NetworkGPU-Accelerated Barrier-Rate Guided MPPI Control for Tractor-Trailer Systems
Articulated vehicles such as tractor-trailers, yard trucks, and similar platforms must often reverse and maneuver in cluttered spaces where pedestrians are present. We present how Barrier-Rate guided Model Predictive Pat…
MP-MPPI: A Motion Primitive Guided Sampling-Based Optimizer for Model Predictive Control
This paper proposes a novel method that extends the Model Predictive Path Integral (MPPI) method with motion primitives for additional structured sampling, which enhances the convergence towards a globally optimal soluti…
Stochastic modeling of cyclic cancer treatments under common noise
Path integral control is an effective method in cancer drug treatment, providing a structured approach to handle the complexities and unpredictability of tumor behavior. Utilizing mathematical principles from physics, th…
DiversityEfficient Implementation of Reinforcement Learning over Homomorphic Encryption
We investigate encrypted control policy synthesis over the cloud. While encrypted control implementations have been studied previously, we focus on the less explored paradigm of privacy-preserving control synthesis, whic…
Privacy Preservingreinforcement-learningReinforcement LearningReinforcement Learning (RL)