paper-with-me

홈 › Papers

A Red Teaming Roadmap Towards System-Level Safety

2025-05-30 · Zifan Wang, Christina Q. Knight, Jeremy Kritz, Willow E. Primack, Julian Michael

Large Language Model (LLM) safeguards, which implement request refusals, have become a widely adopted mitigation strategy against misuse. At the intersection of adversarial machine learning and AI safety, safeguard red teaming has effectively identified critical vulnerabilities in state-of-the-art refusal-trained LLMs. However, in our view the many conference submissions on LLM red teaming do not, in aggregate, prioritize the right research problems. First, testing against clear product safety specifications should take a higher priority than abstract social biases or ethical principles. Second, red teaming should prioritize realistic threat models that represent the expanding risk landscape and what real attackers might do. Finally, we contend that system-level safety is a necessary step to move red teaming research forward, as AI models present new threats as well as affordances for threat mitigation (e.g., detection and banning of malicious users) once placed in a deployment context. Adopting these priorities will be necessary in order for red teaming research to adequately address the slate of new threats that rapid AI advances present today and will present in the very near future.

📄 PDF Abstract BibTeX arXiv:2506.05376

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language ModelRed Teaming

Similar Papers 제목 키워드 기반

Lessons From Red Teaming 100 Generative AI Products

2025-01-13 · Blake Bullwinkel, Amanda Minnich, Shiven Chawla, Gary Lopez 외

In recent years, AI red teaming has emerged as a practice for probing the safety and security of generative AI systems. Due to the nascency of the field, there are many open questions about how red teaming operations sho…

BenchmarkingRed Teaming

X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents

2025-04-15 · Salman Rahman, Liwei Jiang, James Shiffer, Genglin Liu 외

Multi-turn interactions with language models (LMs) pose critical safety risks, as harmful intent can be strategically spread across exchanges. Yet, the vast majority of prior work has focused on single-turn safety, while…

DiversityRed TeamingSafety Alignment

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems

2025-06-16 · Junfeng Fang, Zijun Yao, Ruipeng Wang, Haokai Ma 외

The development of large language models (LLMs) has entered in a experience-driven era, flagged by the emergence of environment feedback-driven learning via reinforcement learning and tool-using agents. This encourages t…

PositionRed Teaming

WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

2024-06-26 · Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger 외

We introduce WildTeaming, an automatic LLM safety red-teaming framework that mines in-the-wild user-chatbot interactions to discover 5.7K unique clusters of novel jailbreak tactics, and then composes multiple tactics for…

ChatbotRed Teaming

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI

2024-12-08 · Chao Yang, Chaochao Lu, Yingchun Wang, BoWen Zhou

Ensuring Artificial General Intelligence (AGI) reliably avoids harmful behaviors is a critical challenge, especially for systems with high autonomy or in safety-critical domains. Despite various safety assurance proposal…

Decision Making