Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
Prompt injection (both direct and indirect) and jailbreaking are now recognized as significant issues for large language models (LLMs), particularly due to their potential for harm in application-integrated contexts. This extended abstract explores a novel approach to protecting LLMs from such attacks, termed "soft begging." This method involves training soft prompts to counteract the effects of corrupted prompts on the LLM's output. We provide an overview of prompt injections and jailbreaking, introduce the theoretical basis of the "soft begging" technique, and discuss an evaluation of its effectiveness.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
An Eye for a Treat: Human Gazing Modulates Begging by Free-ranging Dogs
Interspecific communication plays a critical role in mediating human-animal interactions, particularly in contexts involving access to anthropogenic resources. This study investigates the influence of human gazing on the…
Don’t sweat the small stuff, classify the rest: Sample Shielding to protect text classifiers against adversarial attacks
Deep learning (DL) is being used extensively for text classification. However, researchers have demonstrated the vulnerability of such classifiers to adversarial attacks. Attackers modify the text in a way which misleads…
text-classificationText ClassificationDon't sweat the small stuff, classify the rest: Sample Shielding to protect text classifiers against adversarial attacks
Deep learning (DL) is being used extensively for text classification. However, researchers have demonstrated the vulnerability of such classifiers to adversarial attacks. Attackers modify the text in a way which misleads…
text-classificationText ClassificationScheduling Distributed Flexible Assembly Lines using Safe Reinforcement Learning with Soft Shielding
Highly automated assembly lines enable significant productivity gains in the manufacturing industry, particularly in mass production condition. Nonetheless, challenges persist in job scheduling for make-to-job and mass c…
Safe Reinforcement LearningSchedulingSafe Reinforcement Learning in Black-Box Environments via Adaptive Shielding
Empowering safe exploration of reinforcement learning (RL) agents during training is a critical challenge towards their deployment in many real-world scenarios. When prior knowledge of the domain or task is unavailable, …
reinforcement-learningReinforcement Learning (RL)Safe ExplorationSafe Reinforcement Learning