paper-with-me

Papers

Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks

2022-07-20 · Tim Franzmeyer, Stephen Mcaleer, João F. Henriques, Jakob N. Foerster, Philip H. S. Torr, Adel Bibi, Christian Schroeder de Witt

Autonomous agents deployed in the real world need to be robust against adversarial attacks on sensory inputs. Robustifying agent policies requires anticipating the strongest attacks possible. We demonstrate that existing observation-space attacks on reinforcement learning agents have a common weakness: while effective, their lack of information-theoretic detectability constraints makes them detectable using automated means or human inspection. Detectability is undesirable to adversaries as it may trigger security escalations. We introduce {\epsilon}-illusory, a novel form of adversarial attack on sequential decision-makers that is both effective and of {\epsilon}-bounded statistical detectability. We propose a novel dual ascent algorithm to learn such attacks end-to-end. Compared to existing attacks, we empirically find {\epsilon}-illusory to be significantly harder to detect with automated methods, and a small study with human participants (IRB approval under reference R84123/RE001) suggests they are similarly harder to detect for humans. Our findings suggest the need for better anomaly detectors, as well as effective hardware- and system-level defenses. The project website can be found at https://tinyurl.com/illusory-attacks.

📄 PDF Abstract BibTeX arXiv:2207.10170

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial AttackAdversarial Robustness

Similar Papers 제목 키워드 기반

Characterizing Detectability in 3DGS Poisoning: A Stage-wise Benchmark

2026-06-02 · Quoc-Anh Bui-Huynh, Thanh Duc Ngo, Xue Geng, Kaixin Xu 외 arxiv

3D Gaussian Splatting (3DGS) has rapidly emerged as a leading representation for real-time novel view synthesis, but recent work shows it is vulnerable to diverse poisoning attacks, including illusory object injection, c…

Novel View Synthesis

Balancing detectability and performance of attacks on the control channel of Markov Decision Processes

2021-09-15 · Alessio Russo, Alexandre Proutiere

We investigate the problem of designing optimal stealthy poisoning attacks on the control channel of Markov decision processes (MDPs). This research is motivated by the recent interest of the research community for adver…

Reinforcement Learning (RL)

Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models

2025-09-24 · Zhifang Zhang, Jiahan Zhang, Shengjie Zhou, Qi Wei 외 arxiv

Multimodal pre-trained models (e.g., ImageBind), which align distinct data modalities into a shared embedding space, have shown remarkable success across downstream tasks. However, their increasing adoption raises seriou…

Anomaly Detection

The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?

2024-10-02 · Ching-Chia Kao, Chia-Mu Yu, Chun-Shien Lu, Chu-Song Chen

Vision-Language Models (VLMs) have achieved remarkable performance on a variety of tasks, yet they remain vulnerable to jailbreak attacks that compromise safety and reliability. In this paper, we provide an information-t…

Privacy-Preserving Stealthy Attack Detection in Multi-Agent Control Systems

2021-09-28 · Mohammad Bahrami, Hamidreza Jafarnejadsani

This paper develops a glocal (global-local) attack detection framework to detect stealthy cyber-physical attacks, namely covert attack and zero-dynamics attack, against a class of multi-agent control systems seeking aver…

Privacy Preserving