paper-with-me

홈 › Papers

Multilingual Hidden Prompt Injection Attacks on LLM-Based Academic Reviewing

2025-12-29 · Panagiotis Theocharopoulos, Ajinkya Kulkarni, Mathew Magimai. -Doss arxiv

Large language models (LLMs) are increasingly considered for use in high-impact workflows, including academic peer review. However, LLMs are vulnerable to document-level hidden prompt injection attacks. In this work, we construct a dataset of approximately 500 real academic papers accepted to ICML and evaluate the effect of embedding hidden adversarial prompts within these documents. Each paper is injected with semantically equivalent instructions in four different languages and reviewed using an LLM. We find that prompt injection induces substantial changes in review scores and accept/reject decisions for English, Japanese, and Chinese injections, while Arabic injections produce little to no effect. These results highlight the susceptibility of LLM-based reviewing systems to document-level prompt injection and reveal notable differences in vulnerability across languages.

📄 PDF Abstract BibTeX arXiv:2512.23684

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening

2026-05-27 · Mohan Zhang, Yuqi Jia, Zhen Tan, Steven Jiang 외 arxiv

LLMs are vulnerable to prompt injection attacks. However, this vulnerability has been primarily demonstrated conceptually in academic studies or through a few anecdotal case studies. Its prevalence and impact in real-wor…

SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

2026-04-29 · Yuan Xin, Yixuan Weng, Minjun Zhu, Ying Ling 외 arxiv

As Large Language Models (LLMs) are increasingly integrated into academic peer review, their vulnerability to adversarial hidden prompts, i.e., adversarial instructions embedded in submissions to manipulate outcomes, pos…

LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

2026-05-25 · Lingyao Li, Junjie Xiong, Changjia Zhu, Runlong Yu 외 arxiv

Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adversarial attacks remain poorly understood. We present a systematic be…

Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection

2024-04-07 · Zhilong Wang, Yebo Cao, Peng Liu

Jailbreak attacks on Language Model Models (LLMs) entail crafting prompts aimed at exploiting the models to generate malicious content. Existing jailbreak attacks can successfully deceive the LLMs, however they cannot de…

Language ModelingLanguage Modelling

Applying Pre-trained Multilingual BERT in Embeddings for Improved Malicious Prompt Injection Attacks Detection

2024-09-20 · Md Abdur Rahman, Hossain Shahriar, Fan Wu, Alfredo Cuzzocrea

Large language models (LLMs) are renowned for their exceptional capabilities, and applying to a wide range of applications. However, this widespread use brings significant vulnerabilities. Also, it is well observed that …

Binary Classification