paper-with-me

홈 › Papers

Generating Out of Distribution Adversarial Attack using Latent Space Poisoning

2020-12-09 · Ujjwal Upadhyay, Prerana Mukherjee

Traditional adversarial attacks rely upon the perturbations generated by gradients from the network which are generally safeguarded by gradient guided search to provide an adversarial counterpart to the network. In this paper, we propose a novel mechanism of generating adversarial examples where the actual image is not corrupted rather its latent space representation is utilized to tamper with the inherent structure of the image while maintaining the perceptual quality intact and to act as legitimate data samples. As opposed to gradient-based attacks, the latent space poisoning exploits the inclination of classifiers to model the independent and identical distribution of the training dataset and tricks it by producing out of distribution samples. We train a disentangled variational autoencoder (beta-VAE) to model the data in latent space and then we add noise perturbations using a class-conditioned distribution function to the latent space under the constraint that it is misclassified to the target label. Our empirical results on MNIST, SVHN, and CelebA dataset validate that the generated adversarial examples can easily fool robust l_0, l_2, l_inf norm classifiers designed using provably robust defense mechanisms.

📄 PDF Abstract BibTeX arXiv:2012.05027

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Attack

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Generating Realistic Adversarial Examples for Business Processes using Variational Autoencoders

2024-11-21 · Alexander Stevens, Jari Peeperkorn, Johannes De Smedt, Jochen De Weerdt

In predictive process monitoring, predictive models are vulnerable to adversarial attacks, where input perturbations can lead to incorrect predictions. Unlike in computer vision, where these perturbations are designed to…

Predictive Process Monitoring

Generating Adversarial Attacks in the Latent Space

2023-04-10 · Nitish Shukla, Sudipta Banerjee

Adversarial attacks in the input (pixel) space typically incorporate noise margins such as $L_1$ or $L_{\infty}$-norm to produce imperceptibly perturbed data that confound deep learning networks. Such noise margins confi…

Adversarial AttackGenerative Adversarial Network

LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs

2025-05-16 · Ran Li, Hao Wang, Chengzhi Mao

Efficient red-teaming method to uncover vulnerabilities in Large Language Models (LLMs) is crucial. While recent attacks often use LLMs as optimizers, the discrete language space make gradient-based methods struggle. We …

Red Teaming

A look at adversarial attacks on radio waveforms from discrete latent space

2025-06-11 · Attanasia Garuso, Silvija Kokalj-Filipovic, Yagna Kaasaragadda

Having designed a VQVAE that maps digital radio waveforms into discrete latent space, and yields a perfectly classifiable reconstruction of the original data, we here analyze the attack suppressing properties of VQVAE wh…

Adversarial Attack

An h-space Based Adversarial Attack for Protection Against Few-shot Personalization

2025-07-23 · Xide Xu, Sandesh Kamath, Muhammad Atif Butt, Bogdan Raducanu arxiv

The versatility of diffusion models in generating customized images from few samples raises significant privacy concerns, particularly regarding unauthorized modifications of private content. This concerning issue has re…

Adversarial AttackImage Generation