paper-with-me

홈 › Papers

A test suite of prompt injection attacks for LLM-based machine translation

2024-10-07 · Antonio Valerio Miceli-Barone, Zhifan Sun

LLM-based NLP systems typically work by embedding their input data into prompt templates which contain instructions and/or in-context examples, creating queries which are submitted to a LLM, and then parsing the LLM response in order to generate the system outputs. Prompt Injection Attacks (PIAs) are a type of subversion of these systems where a malicious user crafts special inputs which interfere with the prompt templates, causing the LLM to respond in ways unintended by the system designer. Recently, Sun and Miceli-Barone proposed a class of PIAs against LLM-based machine translation. Specifically, the task is to translate questions from the TruthfulQA test suite, where an adversarial prompt is prepended to the questions, instructing the system to ignore the translation instruction and answer the questions instead. In this test suite, we extend this approach to all the language pairs of the WMT 2024 General Machine Translation task. Moreover, we include additional attack formats in addition to the one originally studied.

📄 PDF Abstract BibTeX arXiv:2410.05047

Code (1)

Avmb/adversarial_MT_prompt_injection 공식 구현 pytorch

Tasks

Machine TranslationTranslationTruthfulQA

Similar Papers 제목 키워드 기반

AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

2024-06-19 · Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner 외

AI agents aim to solve complex tasks by combining text-based reasoning with external tool calls. Unfortunately, AI agents are vulnerable to prompt injection attacks where data returned by external tools hijacks the agent…

BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents

2025-11-25 · Kaiyuan Zhang, Mark Tenenholtz, Kyle Polley, Jerry Ma 외 arxiv

The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web application threat models. Prior work has identified prompt injection as a new attack…

PROMPTFUZZ: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs

2024-09-23 · Jiahao Yu, Yangguang Shao, Hanwen Miao, Junzheng Shi

Large Language Models (LLMs) have gained widespread use in various applications due to their powerful capability to generate human-like text. However, prompt injection attacks, which involve overwriting a model's origina…

Automatic and Universal Prompt Injection Attacks against Large Language Models

2024-03-07 · Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang 외

Large Language Models (LLMs) excel in processing and generating human language, powered by their ability to interpret and follow instructions. However, their capabilities can be exploited through prompt injection attacks…

MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs

2026-02-06 · Junhyeok Lee, Han Jang, Kyu Sung Choi arxiv

Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems are increasingly integrated into clinical workflows; however, prompt injection attacks can steer these systems toward clinically unsafe or mis…