paper-with-me

Papers

Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning

2024-07-03 · Simon Ostermann, Kevin Baum, Christoph Endres, Julia Masloh, Patrick Schramowski

Prompt injection (both direct and indirect) and jailbreaking are now recognized as significant issues for large language models (LLMs), particularly due to their potential for harm in application-integrated contexts. This extended abstract explores a novel approach to protecting LLMs from such attacks, termed "soft begging." This method involves training soft prompts to counteract the effects of corrupted prompts on the LLM's output. We provide an overview of prompt injections and jailbreaking, introduce the theoretical basis of the "soft begging" technique, and discuss an evaluation of its effectiveness.

📄 PDF Abstract BibTeX arXiv:2407.03391

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Eye for a Treat: Human Gazing Modulates Begging by Free-ranging Dogs

2025-05-23 · Sourabh Biswas, Srijaya Nandi, Tuhin Subhra Pal, Aesha Lahiri 외

Interspecific communication plays a critical role in mediating human-animal interactions, particularly in contexts involving access to anthropogenic resources. This study investigates the influence of human gazing on the…

Don’t sweat the small stuff, classify the rest: Sample Shielding to protect text classifiers against adversarial attacks

2022-07-01 · NAACL 2022 7 · Jonathan Rusert, Padmini Srinivasan

Deep learning (DL) is being used extensively for text classification. However, researchers have demonstrated the vulnerability of such classifiers to adversarial attacks. Attackers modify the text in a way which misleads…

text-classificationText Classification

Don't sweat the small stuff, classify the rest: Sample Shielding to protect text classifiers against adversarial attacks

2022-05-03 · Jonathan Rusert, Padmini Srinivasan

Deep learning (DL) is being used extensively for text classification. However, researchers have demonstrated the vulnerability of such classifiers to adversarial attacks. Attackers modify the text in a way which misleads…

text-classificationText Classification

Scheduling Distributed Flexible Assembly Lines using Safe Reinforcement Learning with Soft Shielding

2023-11-21 · Lele Li, Liyong Lin

Highly automated assembly lines enable significant productivity gains in the manufacturing industry, particularly in mass production condition. Nonetheless, challenges persist in job scheduling for make-to-job and mass c…

Safe Reinforcement LearningScheduling

Safe Reinforcement Learning in Black-Box Environments via Adaptive Shielding

2024-05-28 · Daniel Bethell, Simos Gerasimou, Radu Calinescu, Calum Imrie

Empowering safe exploration of reinforcement learning (RL) agents during training is a critical challenge towards their deployment in many real-world scenarios. When prior knowledge of the domain or task is unavailable, …

reinforcement-learningReinforcement Learning (RL)Safe ExplorationSafe Reinforcement Learning