paper-with-me

홈 › Papers

Black Box Adversarial Prompting for Foundation Models

2023-02-08 · Natalie Maus, Patrick Chao, Eric Wong, Jacob Gardner

Prompting interfaces allow users to quickly adjust the output of generative models in both vision and language. However, small changes and design choices in the prompt can lead to significant differences in the output. In this work, we develop a black-box framework for generating adversarial prompts for unstructured image and text generation. These prompts, which can be standalone or prepended to benign prompts, induce specific behaviors into the generative process, such as generating images of a particular object or generating high perplexity text.

📄 PDF Abstract BibTeX arXiv:2302.04237

Code (1)

debugml/adversarial_prompting 공식 구현 pytorch

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models

2026-05-25 · Arian Komaei Koma, Seyed Amir Kasaei, AmirMahdi Sadeghzadeh, Mohammad Hossein Rohban arxiv

Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been introduced to make the model generate such unlearned concepts. These …

Robust Adaptation of Foundation Models with Black-Box Visual Prompting

2024-07-04 · Changdae Oh, Gyeongdeok Seo, Geunyoung Jung, Zhi-Qi Cheng 외

With the surge of large-scale pre-trained models (PTMs), adapting these models to numerous downstream tasks becomes a crucial problem. Consequently, parameter-efficient transfer learning (PETL) of large models has graspe…

Transfer LearningVisual Prompting

JAB: Joint Adversarial Prompting and Belief Augmentation

2023-11-16 · Ninareh Mehrabi, Palash Goyal, Anil Ramakrishna, Jwala Dhamala 외

With the recent surge of language models in different applications, attention to safety and robustness of these models has gained significant importance. Here we introduce a joint framework in which we simultaneously pro…

Red Teaming

Scaling Laws for Black box Adversarial Attacks

2024-11-25 · Chuan Liu, Huanran Chen, Yichi Zhang, Yinpeng Dong 외

Adversarial examples usually exhibit good cross-model transferability, enabling attacks on black-box models with limited information about their architectures and parameters, which are highly threatening in commercial bl…

Adversarial Attack

Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence

2023-04-10 · Hanbin Hong, Xinyu Zhang, Binghui Wang, Zhongjie Ba 외

Black-box adversarial attacks have demonstrated strong potential to compromise machine learning models by iteratively querying the target model or leveraging transferability from a local surrogate model. Recently, such a…

Benchmarkingspeech-recognitionSpeech Recognition