paper-with-me

홈 › Papers

Adversarial attacks hidden in plain sight

2019-02-25 · Jan Philip Göpfert, André Artelt, Heiko Wersing, Barbara Hammer

Convolutional neural networks have been used to achieve a string of successes during recent years, but their lack of interpretability remains a serious issue. Adversarial examples are designed to deliberately fool neural networks into making any desired incorrect classification, potentially with very high certainty. Several defensive approaches increase robustness against adversarial attacks, demanding attacks of greater magnitude, which lead to visible artifacts. By considering human visual perception, we compose a technique that allows to hide such adversarial attacks in regions of high complexity, such that they are imperceptible even to an astute observer. We carry out a user study on classifying adversarially modified images to validate the perceptual quality of our approach and find significant evidence for its concealment with regards to human visual perception.

📄 PDF Abstract BibTeX arXiv:1902.09286

Code (1)

jangop/entropy-based-adversarials pytorch

Tasks

General Classification

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음

Similar Papers 제목 키워드 기반

Why do universal adversarial attacks work on large language models?: Geometry might be the answer

2023-09-01 · Varshini Subhash, Anna Bialas, Weiwei Pan, Finale Doshi-Velez

Transformer based large language models with emergent capabilities are becoming increasingly ubiquitous in society. However, the task of understanding and interpreting their internal workings, in the context of adversari…

Dimensionality Reduction

Hidden in Plain Sight: Undetectable Adversarial Bias Attacks on Vulnerable Patient Populations

2024-02-08 · Pranav Kulkarni, Andrew Chan, Nithya Navarathna, Skylar Chan 외

The proliferation of artificial intelligence (AI) in radiology has shed light on the risk of deep learning (DL) models exacerbating clinical biases towards vulnerable patient populations. While prior literature has focus…

Robust Design of Deep Neural Networks against Adversarial Attacks based on Lyapunov Theory

2019-11-12 · CVPR 2020 6 · Arash Rahnama, Andre T. Nguyen, Edward Raff

Deep neural networks (DNNs) are vulnerable to subtle adversarial perturbations applied to the input. These adversarial perturbations, though imperceptible, can easily mislead the DNN. In this work, we take a control theo…

Robust Design

Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH

2026-06-12 · Vikhyath Kothamasu, Virginia Smith, Chhavi Yadav arxiv

LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-world. A key emerging threat is Decomposition Attacks \cite{glukhov2024breach, jones2…

Hiding-in-Plain-Sight (HiPS) Attack on CLIP for Targetted Object Removal from Images

2024-10-16 · Arka Daw, Megan Hong-Thanh Chung, Maria Mahbub, Amir Sadovnik

Machine learning models are known to be vulnerable to adversarial attacks, but traditional attacks have mostly focused on single-modalities. With the rise of large multi-modal models (LMMs) like CLIP, which combine visio…

Image CaptioningObject