Balancing detectability and performance of attacks on the control channel of Markov Decision Processes
We investigate the problem of designing optimal stealthy poisoning attacks on the control channel of Markov decision processes (MDPs). This research is motivated by the recent interest of the research community for adversarial and poisoning attacks applied to MDPs, and reinforcement learning (RL) methods. The policies resulting from these methods have been shown to be vulnerable to attacks perturbing the observations of the decision-maker. In such an attack, drawing inspiration from adversarial examples used in supervised learning, the amplitude of the adversarial perturbation is limited according to some norm, with the hope that this constraint will make the attack imperceptible. However, such constraints do not grant any level of undetectability and do not take into account the dynamic nature of the underlying Markov process. In this paper, we propose a new attack formulation, based on information-theoretical quantities, that considers the objective of minimizing the detectability of the attack as well as the performance of the controlled process. We analyze the trade-off between the efficiency of the attack and its detectability. We conclude with examples and numerical simulations illustrating this trade-off.
Code (1)
Tasks
Reinforcement Learning (RL)Similar Papers 제목 키워드 기반
Cyber-Attack Detection in Discrete Nonlinear Multi-Agent Systems Using Neural Networks
This paper proposes a distributed cyber-attack detection method in communication channels for a class of discrete, nonlinear, heterogeneous, multi-agent systems that are controlled by our proposed formation-based control…
Cyber Attack DetectionIllusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks
Autonomous agents deployed in the real world need to be robust against adversarial attacks on sensory inputs. Robustifying agent policies requires anticipating the strongest attacks possible. We demonstrate that existing…
Adversarial AttackAdversarial RobustnessRobust LLM Watermarking with Minimal Semantic Distortion for IP Protection
Proprietary large language models (LLMs) face risks of intellectual property (IP) violation, as adversaries can replicate an LLM by collecting input-output pairs to train a surrogate model, causing financial setbacks. Wa…
Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models
Multimodal pre-trained models (e.g., ImageBind), which align distinct data modalities into a shared embedding space, have shown remarkable success across downstream tasks. However, their increasing adoption raises seriou…
Anomaly DetectionEnhancing sensor attack detection in supervisory control systems modeled by probabilistic automata
Sensor attacks compromise the reliability of cyber-physical systems (CPSs) by altering sensor outputs with the objective of leading the system to unsafe system states. This paper studies a probabilistic intrusion detecti…
Intrusion Detection