paper-with-me

홈 › Papers

Deceiving Google's Perspective API Built for Detecting Toxic Comments

2017-02-27 · Hossein Hosseini, Sreeram Kannan, Baosen Zhang, Radha Poovendran

Social media platforms provide an environment where people can freely engage in discussions. Unfortunately, they also enable several problems, such as online harassment. Recently, Google and Jigsaw started a project called Perspective, which uses machine learning to automatically detect toxic language. A demonstration website has been also launched, which allows anyone to type a phrase in the interface and instantaneously see the toxicity score [1]. In this paper, we propose an attack on the Perspective toxic detection system based on the adversarial examples. We show that an adversary can subtly modify a highly toxic phrase in a way that the system assigns significantly lower toxicity score to it. We apply the attack on the sample phrases provided in the Perspective website and show that we can consistently reduce the toxicity scores to the level of the non-toxic phrases. The existence of such adversarial examples is very harmful for toxic detection systems and seriously undermines their usability.

📄 PDF Abstract BibTeX arXiv:1702.08138

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Jigsaw Jigsaw is a self-supervision approach that relies on jigsaw-like puzzles as the pretext task in order to learn image representations.

Similar Papers 제목 키워드 기반

Reading Between the Demographic Lines: Resolving Sources of Bias in Toxicity Classifiers

2020-06-29 · Elizabeth Reichert, Helen Qiu, Jasmine Bayrooti

The censorship of toxic comments is often left to the judgment of imperfect models. Perspective API, a creation of Google technology incubator Jigsaw, is perhaps the most widely used toxicity classifier in industry; the …

Critical Perspectives: A Benchmark Revealing Pitfalls in PerspectiveAPI

2023-01-05 · Lorena Piedras, Lucas Rosenblatt, Julia Wilkins

Detecting "toxic" language in internet content is a pressing social and technical challenge. In this work, we focus on PERSPECTIVE from Jigsaw, a state-of-the-art tool that promises to score the "toxicity" of text, with …

Binary Classification

How toxic is antisemitism? Potentials and limitations of automated toxicity scoring for antisemitic online content

2023-10-05 · Helena Mihaljević, Elisabeth Steffen

The Perspective API, a popular text toxicity assessment service by Google and Jigsaw, has found wide adoption in several application areas, notably content moderation, monitoring, and social media research. We examine it…

ConvAI at SemEval-2019 Task 6: Offensive Language Identification and Categorization with Perspective and BERT

2019-06-01 · SEMEVAL 2019 6 · John Pavlopoulos, Nithum Thain, Lucas Dixon, Ion Androutsopoulos

This paper presents the application of two strong baseline systems for toxicity detection and evaluates their performance in identifying and categorizing offensive language in social media. PERSPECTIVE is an API, that se…

Language Identification

SafeSpace: An Integrated Web Application for Digital Safety and Emotional Well-being

2025-08-22 · Kayenat Fatmi, Mohammad Abbas arxiv

In the digital era, individuals are increasingly exposed to online harms such as toxicity, manipulation, and grooming, which often pose emotional and safety risks. Existing systems for detecting abusive content or issuin…