paper-with-me

홈 › Papers

All You Need is "Love": Evading Hate-speech Detection

2018-08-28 · Tommi Gröndahl, Luca Pajola, Mika Juuti, Mauro Conti, N. Asokan

With the spread of social networks and their unfortunate use for hate speech, automatic detection of the latter has become a pressing problem. In this paper, we reproduce seven state-of-the-art hate speech detection models from prior work, and show that they perform well only when tested on the same type of data they were trained on. Based on these results, we argue that for successful hate speech detection, model architecture is less important than the type of data and labeling criteria. We further show that all proposed detection techniques are brittle against adversaries who can (automatically) insert typos, change word boundaries or add innocuous words to the original hate speech. A combination of these methods is also effective against Google Perspective -- a cutting-edge solution from industry. Our experiments demonstrate that adversarial training does not completely mitigate the attacks, and using character-level features makes the models systematically more attack-resistant than using word-level features.

📄 PDF Abstract BibTeX arXiv:1808.09115

Code (0)

등록된 구현이 없습니다.

Tasks

AllHate Speech Detection

Similar Papers 제목 키워드 기반

All You Need is "Leet": Evading Hate-speech Detection AI

2025-05-22 · Sampanna Yashwant Kahu, Naman Ahuja

Social media and online forums are increasingly becoming popular. Unfortunately, these platforms are being used for spreading hate speech. In this paper, we design black-box techniques to protect users from hate-speech o…

AllHate Speech Detection

INF-HatEval at SemEval-2019 Task 5: Convolutional Neural Networks for Hate Speech Detection Against Women and Immigrants on Twitter

2019-06-01 · SEMEVAL 2019 6 · Alison Ribeiro, N{\'a}dia Silva

In this paper, we describe our approach to detect hate speech against women and immigrants on Twitter in a multilingual context, English and Spanish. This challenge was proposed by the SemEval-2019 Task 5, where particip…

Hate Speech DetectionWord Embeddings

Hate speech detection using static BERT embeddings

2021-06-29 · Gaurav Rajput, Narinder Singh Punn, Sanjay Kumar Sonbhadra, Sonali Agarwal

With increasing popularity of social media platforms hate speech is emerging as a major concern, where it expresses abusive speech that targets specific group characteristics, such as gender, religion or ethnicity to spr…

Hate Speech DetectionSpecificityWord Embeddings

A study of text representations in Hate Speech Detection

2021-02-08 · Chrysoula Themeli, George Giannakopoulos, Nikiforos Pittaras

The pervasiveness of the Internet and social media have enabled the rapid and anonymous spread of Hate Speech content on microblogging platforms such as Twitter. Current EU and US legislation against hateful language, in…

Abusive LanguageHate Speech DetectionWord Embeddings

Exploring Stylometric and Emotion-Based Features for Multilingual Cross-Domain Hate Speech Detection

2021-04-01 · EACL (WASSA) 2021 4 · Ilia Markov, Nikola Ljubešić, Darja Fišer, Walter Daelemans

In this paper, we describe experiments designed to evaluate the impact of stylometric and emotion-based features on hate speech detection: the task of classifying textual content into hate or non-hate speech classes. Our…

Hate Speech Detection