Adversarial Attacks on Image Generation With Made-Up Words
Text-guided image generation models can be prompted to generate images using nonce words adversarially designed to robustly evoke specific visual concepts. Two approaches for such generation are introduced: macaronic prompting, which involves designing cryptic hybrid words by concatenating subword units from different languages; and evocative prompting, which involves designing nonce words whose broad morphological features are similar enough to that of existing words to trigger robust visual associations. The two methods can also be combined to generate images associated with more specific visual concepts. The implications of these techniques for the circumvention of existing approaches to content moderation, and particularly the generation of offensive or harmful images, are discussed.
Code (0)
등록된 구현이 없습니다.
Tasks
Image GenerationSimilar Papers 제목 키워드 기반
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
Recent studies show that text-to-image (T2I) models are vulnerable to adversarial attacks, especially with noun perturbations in text prompts. In this study, we investigate the impact of adversarial attacks on different …
Adversarial AttackImage GenerationPOSTAG+2Controlled Caption Generation for Images Through Adversarial Attacks
Deep learning is found to be vulnerable to adversarial examples. However, its adversarial susceptibility in image caption generation is under-explored. We study adversarial examples for vision and language models, which …
Caption GenerationImage CaptioningLanguage ModellingUniversal Adversarial Attacks with Natural Triggers for Text Classification
Recent work has demonstrated the vulnerability of modern text classifiers to universal adversarial attacks, which are input-agnostic sequences of words added to text processed by classifiers. Despite being successful, th…
ClassificationGeneral Classificationtext-classificationText ClassificationExact Adversarial Attack to Image Captioning via Structured Output Learning with Latent Variables
In this work, we study the robustness of a CNN+RNN based image captioning system being subjected to adversarial noises. We propose to fool an image captioning system to generate some targeted partial captions for an imag…
Adversarial AttackImage CaptioningFall Leaf Adversarial Attack on Traffic Sign Classification
Adversarial input image perturbation attacks have emerged as a significant threat to machine learning algorithms, particularly in image classification setting. These attacks involve subtle perturbations to input images t…
Adversarial AttackClassificationEdge Detectionimage-classification+1