Red Teaming
1개 벤치마크 · 논문 355편 · 이 태스크의 논문 보기 →
Benchmarks
SUDO Dataset
Most implemented
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Explore, Establish, Exploit: Red Teaming Language Models from Scratch
SIR: Self-improving Red-teaming for Compute Use Agents
Papers
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations…
Red TeamingSIR: Self-improving Red-teaming for Compute Use Agents
Computer use agents (CUAs) are vision-language models that perceive a screen and act on a real operating system through mouse, keyboard, and terminal, and they are increasingly deployed to automate everyday digital tasks…
Red TeamingPsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies
Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engin…
Reinforcement LearningRed TeamingFrom Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model an…
Red TeamingAgent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming
Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing s…
Reinforcement LearningRed TeamingOpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that …
Red Teaming