paper-with-me

홈 › Papers

Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming

2023-11-10 · Nanna Inie, Jonathan Stray, Leon Derczynski

Engaging in the deliberate generation of abnormal outputs from Large Language Models (LLMs) by attacking them is a novel human activity. This paper presents a thorough exposition of how and why people perform such attacks, defining LLM red-teaming based on extensive and diverse evidence. Using a formal qualitative methodology, we interviewed dozens of practitioners from a broad range of backgrounds, all contributors to this novel work of attempting to cause LLMs to fail. We focused on the research questions of defining LLM red teaming, uncovering the motivations and goals for performing the activity, and characterizing the strategies people use when attacking LLMs. Based on the data, LLM red teaming is defined as a limit-seeking, non-malicious, manual activity, which depends highly on a team-effort and an alchemist mindset. It is highly intrinsically motivated by curiosity, fun, and to some degrees by concerns for various harms of deploying LLMs. We identify a taxonomy of 12 strategies and 35 different techniques of attacking LLMs. These findings are presented as a comprehensive grounded theory of how and why people attack large language models: LLM red teaming.

📄 PDF Abstract BibTeX arXiv:2311.06237

Code (0)

등록된 구현이 없습니다.

Tasks

Red Teaming

Similar Papers 제목 키워드 기반

Measuring Successful Cooperation in Human-AI Teamwork: Development and Validation of the Perceived Cooperativity and Teaming Perception Scales

2026-04-27 · Christiane Attig, Christiane Wiebel-Herboth, Patricia Wollstadt, Tim Schrills 외 arxiv

As human-AI cooperation becomes increasingly prevalent, reliable instruments for assessing the subjective quality of cooperative human-AI interaction are needed. We introduce two theoretically grounded scales: the Percei…

Scene Synthesis from Human Motion

2023-01-04 · Sifan Ye, Yixing Wang, Jiaman Li, Dennis Park 외

Large-scale capture of human motion with diverse, complex scenes, while immensely useful, is often considered prohibitively costly. Meanwhile, human motion alone contains rich information about the scene they reside in a…

2D Semantic Segmentation task 1 (8 classes)3D Semantic Scene CompletionIndoor Scene Synthesis

Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming

2026-01-26 · Alexandra Chouldechova, A. Feder Cooper, Solon Barocas, Abhinav Palia 외 arxiv

We argue that conclusions drawn about relative system safety or attack method efficacy via AI red teaming are often not supported by evidence provided by attack success rate (ASR) comparisons. We show, through conceptual…

Red Teaming

REALM: A Unified Red-Teaming Benchmark for Physical-World VLMs

2026-06-22 · Yifei Zhao, Qian Lou, Mengxin Zheng arxiv

Vision-language models (VLMs) are increasingly used as perception-reasoning backbones for embodied intelligence in safety-critical physical systems, where perception or reasoning errors can lead to unsafe decisions or ac…

Adversarial Robustness

Red Teaming AI Red Teaming

2025-07-07 · Subhabrata Majumdar, Brian Pendleton, Abhishek Gupta arxiv

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that de…

Red Teaming