paper-with-me

홈 › Papers

Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition

2023-10-24 · Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, Jordan Boyd-Graber

Large Language Models (LLMs) are deployed in interactive contexts with direct user engagement, such as chatbots and writing assistants. These deployments are vulnerable to prompt injection and jailbreaking (collectively, prompt hacking), in which models are manipulated to ignore their original instructions and follow potentially malicious ones. Although widely acknowledged as a significant security threat, there is a dearth of large-scale resources and quantitative studies on prompt hacking. To address this lacuna, we launch a global prompt hacking competition, which allows for free-form human input attacks. We elicit 600K+ adversarial prompts against three state-of-the-art LLMs. We describe the dataset, which empirically verifies that current LLMs can indeed be manipulated via prompt hacking. We also present a comprehensive taxonomical ontology of the types of adversarial prompts.

📄 PDF Abstract BibTeX arXiv:2311.16119

Code (3)

promptlabs/hackaprompt 공식 구현
trigaten/learn_prompting 공식 구현
lostoxygen/llm-confidentiality pytorch

Methods 이 논문이 사용한 방법론

Ontology 설명 없음

Similar Papers 제목 키워드 기반

Uni-FinLLM: A Unified Multimodal Large Language Model with Modular Task Heads for Micro-Level Stock Prediction and Macro-Level Systemic Risk Assessment

2026-01-06 · Gongao Zhang, Haijiang Zeng, Lu Jiang arxiv

Financial institutions and regulators require systems that integrate heterogeneous data to assess risks from stock fluctuations to systemic vulnerabilities. Existing approaches often treat these tasks in isolation, faili…

What Makes Systemic Discrimination, "Systemic?" Exposing the Amplifiers of Inequity

2024-03-16 · David B. McMillon

Drawing on work spanning economics, public health, education, sociology, and law, I formalize theoretically what makes systemic discrimination "systemic." Injustices do not occur in isolation, but within a complex system…

Sociology

Vulnerability Webs: Systemic Risk in Software Networks

2024-02-20 · Cornelius Fritz, Co-Pierre Georg, Angelo Mele, Michael Schweinberger

Modern software development is a collaborative effort that re-uses existing code to reduce development and maintenance costs. This practice exposes software to vulnerabilities in the form of undetected bugs in direct and…

Unveiling the Landscape of LLM Deployment in the Wild: An Empirical Study

2025-05-05 · Xinyi Hou, Jiahao Han, Yanjie Zhao, Haoyu Wang

Background: Large language models (LLMs) are increasingly deployed via open-source and commercial frameworks, enabling individuals and organizations to self-host advanced AI capabilities. However, insecure defaults and m…

Red Teaming Large Language Models for Healthcare

2025-05-01 · Vahid Balazadeh, Michael Cooper, David Pellow, Atousa Assadi 외

We present the design process and findings of the pre-conference workshop at the Machine Learning for Healthcare Conference (2024) entitled Red Teaming Large Language Models for Healthcare, which took place on August 15,…

Language ModelingLanguage ModellingLarge Language ModelRed Teaming