paper-with-me

Red Teaming

1개 벤치마크 · 논문 355편 · 이 태스크의 논문 보기 →

Benchmarks

SUDO Dataset

결과 1개

Most implemented

Papers

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

2026-09-09 · Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal 외 arxiv

Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations…

Red Teaming

SIR: Self-improving Red-teaming for Compute Use Agents

2026-08-31 · Chen Xiong, Zhiyuan He, Pin-Yu Chen, Stjepan Picek 외 arxiv

Computer use agents (CUAs) are vision-language models that perceive a screen and act on a real operating system through mouse, keyboard, and terminal, and they are increasingly deployed to automate everyday digital tasks…

Red Teaming

PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

2026-08-24 · Zeyu Feng, Qingyu Wu, Yuzhe Luo, Hua Cheng arxiv

Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engin…

Reinforcement LearningRed Teaming

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

2026-08-10 · Yuanhe Zhang, Weiliu Wang, Jie Ren, Liang Lin 외 hf

Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model an…

Red Teaming

Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming

2026-08-05 · Yanting Wang, Chenlong Yin, Runpeng Geng, Jinyuan Jia hf

Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing s…

Reinforcement LearningRed Teaming

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

2026-08-01 · Yunhao Chen, Xin Wang, Yixu Wang, Yi Liu 외 arxiv

AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through a shared state that …

Red Teaming

전체 355편 보기 →