paper-with-me

홈 › Papers

TaeBench: Improving Quality of Toxic Adversarial Examples

2024-10-08 · Xuan Zhu, Dmitriy Bespalov, Liwen You, Ninad Kulkarni, Yanjun Qi

Toxicity text detectors can be vulnerable to adversarial examples - small perturbations to input text that fool the systems into wrong detection. Existing attack algorithms are time-consuming and often produce invalid or ambiguous adversarial examples, making them less useful for evaluating or improving real-world toxicity content moderators. This paper proposes an annotation pipeline for quality control of generated toxic adversarial examples (TAE). We design model-based automated annotation and human-based quality verification to assess the quality requirements of TAE. Successful TAE should fool a target toxicity model into making benign predictions, be grammatically reasonable, appear natural like human-generated text, and exhibit semantic toxicity. When applying these requirements to more than 20 state-of-the-art (SOTA) TAE attack recipes, we find many invalid samples from a total of 940k raw TAE attack generations. We then utilize the proposed pipeline to filter and curate a high-quality TAE dataset we call TaeBench (of size 264k). Empirically, we demonstrate that TaeBench can effectively transfer-attack SOTA toxicity content moderation models and services. Our experiments also show that TaeBench with adversarial training achieve significant improvements of the robustness of two toxicity detectors.

📄 PDF Abstract BibTeX arXiv:2410.05573

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Building a Robust Toxicity Predictor

2024-04-09 · Dmitriy Bespalov, Sourav Bhabesh, Yi Xiang, Liutong Zhou 외

Recent NLP literature pays little attention to the robustness of toxicity language predictors, while these systems are most likely to be used in adversarial contexts. This paper presents a novel adversarial attack, \text…

Adversarial Attack

Deceiving Google's Perspective API Built for Detecting Toxic Comments

2017-02-27 · Hossein Hosseini, Sreeram Kannan, Baosen Zhang, Radha Poovendran

Social media platforms provide an environment where people can freely engage in discussions. Unfortunately, they also enable several problems, such as online harassment. Recently, Google and Jigsaw started a project call…

Fortifying Toxic Speech Detectors Against Veiled Toxicity

2020-10-07 · EMNLP 2020 11 · Xiaochuang Han, Yulia Tsvetkov

Modern toxic speech detectors are incompetent in recognizing disguised offensive language, such as adversarial attacks that deliberately avoid known toxic lexicons, or manifestations of implicit bias. Building a large an…

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF

2026-04-20 · Yuan Fang, Yiming Luo, Aimin Zhou, Fei Tan arxiv

Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-explored. We propose Reverse Constitutional AI (R-CAI), a framework f…

Reinforcement LearningRed Teaming

ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

2022-03-17 · ACL 2022 5 · Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap 외

Toxic language detection systems often falsely flag text that contains minority group mentions as toxic, as those groups are often the targets of online hate. Such over-reliance on spurious correlations also causes syste…

Hate Speech DetectionLanguage Modelling