paper-with-me

홈 › Papers

Intent Laundering: AI Safety Datasets Are Not What They Seem

2026-02-17 · Shahriar Golchin, Marc Wetter arxiv

We systematically evaluate the quality of widely used adversarial safety datasets from two perspectives: in isolation and in practice. In isolation, we examine how well these datasets reflect real-world adversarial attacks based on three defining properties: being driven by ulterior intent, well-crafted, and out-of-distribution. We find that these datasets overrely on "triggering cues": words or phrases with overt negative/sensitive connotations that are intended to trigger safety mechanisms explicitly, which is unrealistic compared to real-world attacks. In practice, we evaluate whether these datasets genuinely measure safety risks or merely provoke refusals through triggering cues. To explore this, we introduce "intent laundering": a procedure that abstracts away triggering cues from adversarial attacks (data points) while strictly preserving their malicious intent and all relevant details. Our results show that current adversarial safety datasets fail to faithfully represent real-world adversarial behavior due to their overreliance on triggering cues. Once these cues are removed, all previously evaluated "reasonably safe" models become unsafe, including Gemini 3 Pro and Claude Sonnet 3.7/4. Moreover, when intent laundering is adapted as a jailbreaking technique, it consistently achieves high attack success rates, ranging from 90.00% to 100.00%, under fully black-box access. Overall, our findings expose a significant disconnect between how existing datasets evaluate model safety and how real-world adversaries behave.

📄 PDF Abstract BibTeX arXiv:2602.16729

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AMLgentex: Mobilizing Data-Driven Research to Combat Money Laundering

2025-06-03 · Johan Östman, Edvin Callisen, Anton Chen, Kristiina Ausmees 외

Money laundering enables organized crime by allowing illicit funds to enter the legitimate economy. Although trillions of dollars are laundered each year, only a small fraction is ever uncovered. This stems from a range …

Benchmarking

Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense

2026-06-28 · Subhadip Mitra arxiv

Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no defense, static steering, CAST, AlphaSteer, probe-gated) across seven…

Laundering AI Authority with Adversarial Examples

2026-05-05 · Jie Zhang, Pura Peetathawatchai, Florian Tramèr, Avital Shafran arxiv

Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the s…

Adversarial Robustness

Intelligent Anti-Money Laundering Solution Based upon Novel Community Detection in Massive Transaction Networks on Spark

2025-01-08 · Xurui Li, Xiang Cao, Xuetao Qiu, Jintao Zhao 외

Criminals are using every means available to launder the profits from their illegal activities into ostensibly legitimate assets. Meanwhile, most commercial anti-money laundering systems are still rule-based, which canno…

Community Detection

The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies

2025-09-29 · Sam Coggins, Alexander K. Saeri, Katherine A. Daniell, Lorenn P. Ruster 외 arxiv

Prominent AI companies are producing 'safety frameworks' as a type of voluntary self-governance. These statements purport to establish risk thresholds and safety procedures for the development and deployment of highly ca…