paper-with-me

홈 › Papers

Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks

2026-02-16 · Lukas Struppek, Adam Gleave, Kellin Pelrine arxiv

As the capabilities of large language models continue to advance, so does their potential for misuse. While closed-source models typically rely on external defenses, open-weight models must primarily depend on internal safeguards to mitigate harmful behavior. Prior red-teaming research has largely focused on input-based jailbreaking and parameter-level manipulations. However, open-weight models also natively support prefilling, which allows an attacker to predefine initial response tokens before generation begins. Despite its potential, this attack vector has received little systematic attention. We present the largest empirical study to date of prefill attacks, evaluating over 20 existing and novel strategies across multiple model families and state-of-the-art open-weight models. Our results show that prefill attacks are consistently effective against all major contemporary open-weight models, revealing a critical and previously underexplored vulnerability with significant implications for deployment. While certain large reasoning models exhibit some robustness against generic prefilling, they remain vulnerable to tailored, model-specific strategies. Our findings underscore the urgent need for model developers to prioritize defenses against prefill attacks in open-weight LLMs.

📄 PDF Abstract BibTeX arXiv:2602.14689

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities

2026-08-26 · Tianshi Wang, Jingsong Wang, Yafei Huang, Fengling Li 외 arxiv

Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple h…

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

2026-05-26 · Kevin Kuo, Chhavi Yadav, Virginia Smith arxiv

Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rat…

When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection

2025-12-11 · Devanshu Sahoo, Manish Prasad, Vasudev Majhi, Jahnvi Singh 외 arxiv

Driven by surging submission volumes, scientific peer review has catalyzed two parallel trends: individual over-reliance on LLMs and institutional AI-powered assessment systems. This study investigates the robustness of …

Matching Ranks Over Probability Yields Truly Deep Safety Alignment

2025-12-05 · Jason Vega, Gagandeep Singh arxiv

A frustratingly easy technique known as the prefilling attack has been shown to effectively circumvent the safety alignment of frontier LLMs by simply prefilling the assistant response with an affirmative prefix before d…

Data Augmentation

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving

2026-05-03 · Zhaoyuan Su, Olatunji Ruwase, Karthik Ganesan, Aurick Qiao 외 arxiv

Production LLM workloads increasingly serve discriminative tasks, such as classification, recommendation, and verification, whose answers are read from the logits of a single prefill pass with no autoregressive decoding.…