paper-with-me

홈 › Papers

Death by a Thousand Prompts: Open Model Vulnerability Analysis

2025-11-05 · Amy Chang, Nicholas Conley, Harish Santhanalakshmi Ganesan, Adam Swanda arxiv

Open-weight models provide researchers and developers with accessible foundations for diverse downstream applications. We tested the safety and security postures of eight open-weight large language models (LLMs) to identify vulnerabilities that may impact subsequent fine-tuning and deployment. Using automated adversarial testing, we measured each model's resilience against single-turn and multi-turn prompt injection and jailbreak attacks. Our findings reveal pervasive vulnerabilities across all tested models, with multi-turn attacks achieving success rates between 25.86\% and 92.78\% -- representing a $2\times$ to $10\times$ increase over single-turn baselines. These results underscore a systemic inability of current open-weight models to maintain safety guardrails across extended interactions. We assess that alignment strategies and lab priorities significantly influence resilience: capability-focused models such as Llama 3.3 and Qwen 3 demonstrate higher multi-turn susceptibility, whereas safety-oriented designs such as Google Gemma 3 exhibit more balanced performance. The analysis concludes that open-weight models, while crucial for innovation, pose tangible operational and ethical risks when deployed without layered security controls. These findings are intended to inform practitioners and developers of the potential risks and the value of professional AI security solutions to mitigate exposure. Addressing multi-turn vulnerabilities is essential to ensure the safe, reliable, and responsible deployment of open-weight LLMs in enterprise and public domains. We recommend adopting a security-first design philosophy and layered protections to ensure resilient deployments of open-weight models.

📄 PDF Abstract BibTeX arXiv:2511.03247

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Automated software vulnerability detection with machine learning

2018-02-14 · Jacob A. Harer, Louis Y. Kim, Rebecca L. Russell, Onur Ozdemir 외

Thousands of security vulnerabilities are discovered in production software each year, either reported publicly to the Common Vulnerabilities and Exposures database or discovered internally in proprietary code. Vulnerabi…

BIG-bench Machine LearningVulnerability Detection

Evaluating LLMs for Real-World Web Vulnerability Detection

2026-06-19 · Sebastian Neef, Luca Jungnickel, Antonio Benjamin Buchholz, Valene Spence 외 arxiv

Large Language Models (LLMs) have emerged as a promising tool for automated vulnerability detection, yet their effectiveness on web-specific vulnerabilities remains to be explored. This work benchmarks six frontier (Clau…

Vulnerability Detection

A First Look at GPT Apps: Landscape and Vulnerability

2024-02-23 · Zejun Zhang, Li Zhang, Xin Yuan, Anlan Zhang 외

Following OpenAI's introduction of GPTs, a surge in GPT apps has led to the launch of dedicated LLM app stores. Nevertheless, given its debut, there is a lack of sufficient understanding of this new ecosystem. To fill th…

Arbiter: Detecting Interference in LLM Agent System Prompts

2026-03-09 · Tony Mason arxiv

System prompts for LLM-based coding agents are software artifacts that govern agent behavior, yet lack the testing infrastructure applied to conventional software. We present Arbiter, a framework combining formal evaluat…

Jailbreaking as a Reward Misspecification Problem

2024-06-20 · Zhihui Xie, Jiahui Gao, Lei LI, Zhenguo Li 외

The widespread adoption of large language models (LLMs) has raised concerns about their safety and reliability, particularly regarding their vulnerability to adversarial attacks. In this paper, we propose a novel perspec…

Red Teaming