paper-with-me

홈 › Papers

A Tale of Evil Twins: Adversarial Inputs versus Poisoned Models

2019-11-05 · Ren Pang, Hua Shen, Xinyang Zhang, Shouling Ji, Yevgeniy Vorobeychik, Xiapu Luo, Alex Liu, Ting Wang

Despite their tremendous success in a range of domains, deep learning systems are inherently susceptible to two types of manipulations: adversarial inputs -- maliciously crafted samples that deceive target deep neural network (DNN) models, and poisoned models -- adversely forged DNNs that misbehave on pre-defined inputs. While prior work has intensively studied the two attack vectors in parallel, there is still a lack of understanding about their fundamental connections: what are the dynamic interactions between the two attack vectors? what are the implications of such interactions for optimizing existing attacks? what are the potential countermeasures against the enhanced attacks? Answering these key questions is crucial for assessing and mitigating the holistic vulnerabilities of DNNs deployed in realistic settings. Here we take a solid step towards this goal by conducting the first systematic study of the two attack vectors within a unified framework. Specifically, (i) we develop a new attack model that jointly optimizes adversarial inputs and poisoned models; (ii) with both analytical and empirical evidence, we reveal that there exist intriguing "mutual reinforcement" effects between the two attack vectors -- leveraging one vector significantly amplifies the effectiveness of the other; (iii) we demonstrate that such effects enable a large design spectrum for the adversary to enhance the existing attacks that exploit both vectors (e.g., backdoor attacks), such as maximizing the attack evasiveness with respect to various detection methods; (iv) finally, we discuss potential countermeasures against such optimized attacks and their technical challenges, pointing to several promising research directions.

📄 PDF Abstract BibTeX arXiv:1911.01559

Code (1)

ain-soph/trojanzoo 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Prompts have evil twins

2023-11-13 · Rimon Melamed, Lucas H. McCabe, Tanay Wakhare, Yejin Kim 외

We discover that many natural-language prompts can be replaced by corresponding prompts that are unintelligible to humans but that provably elicit similar behavior in language models. We call these prompts "evil twins" b…

Evil twins are not that evil: Qualitative insights into machine-generated prompts

2024-12-11 · Nathanaël Carraz Rakotonirina, Corentin Kervadec, Francesca Franzon, Marco Baroni

It has been widely observed that language models (LMs) respond in predictable ways to algorithmically generated prompts that are seemingly unintelligible. This is both a sign that we lack a full understanding of how LMs …

Planar Juggling of a Devil-Stick using Discrete VHCs

2025-09-09 · Aakash Khandelwal, Ranjan Mukherjee arxiv

Planar juggling of a devil-stick using impulsive inputs is addressed using the concept of discrete virtual holonomic constraints (DVHC). The location of the center-of-mass of the devil-stick is specified in terms of its …

Why Don't You Clean Your Glasses? Perception Attacks with Dynamic Optical Perturbations

2023-07-24 · Yi Han, Matthew Chan, Eric Wengrowski, Zhuohuan Li 외

Camera-based autonomous systems that emulate human perception are increasingly being integrated into safety-critical platforms. Consequently, an established body of literature has emerged that explores adversarial attack…

Better the Devil you Know: An Analysis of Evasion Attacks using Out-of-Distribution Adversarial Examples

2019-05-05 · Vikash Sehwag, Arjun Nitin Bhagoji, Liwei Song, Chawin Sitawarin 외

A large body of recent work has investigated the phenomenon of evasion attacks using adversarial examples for deep learning systems, where the addition of norm-bounded perturbations to the test inputs leads to incorrect …

Autonomous DrivingGeneral Classification