Separating Hate Speech and Offensive Language Classes via Adversarial Debiasing
Research to tackle hate speech plaguing online media has made strides in providing solutions, analyzing bias and curating data. A challenging problem is ambiguity between hate speech and offensive language, causing low performance both overall and specifically for the hate speech class. It can be argued that misclassifying actual hate speech content as merely offensive can lead to further harm against targeted groups. In our work, we mitigate this potentially harmful phenomenon by proposing an adversarial debiasing method to separate the two classes. We show that our method works for English, Arabic German and Hindi, plus in a multilingual setting, improving performance over baselines.
Code (1)
Similar Papers 제목 키워드 기반
Afaan Oromo Hate Speech Detection and Classification on Social Media
Hate and offensive speech on social media is targeted to attack an individual or group of community based on protected characteristics such as gender, ethnicity, and religion. Hate and offensive speech on social media is…
ClassificationHate Speech DetectionOverview of the HASOC Subtrack at FIRE 2021: Hate Speech and Offensive Content Identification in English and Indo-Aryan Languages
The widespread of offensive content online such as hate speech poses a growing societal problem. AI tools are necessary for supporting the moderation process at online platforms. For the evaluation of these identificatio…
Binary ClassificationClassificationUsing Transfer-based Language Models to Detect Hateful and Offensive Language Online
Distinguishing hate speech from non-hate offensive language is challenging, as hate speech not always includes offensive slurs and offensive language not always express hate. Here, four deep learners based on the Bidirec…
HateBR: Large expert annotated corpus of Brazilian Instagram comments for abusive language detection
Due to the severity of the social media abusive comments in Brazil, and the lack of research in Portuguese, this paper provides the first large-scale annotated corpus of Brazilian Instagram comments for hate speech and o…
Abusive LanguageBinary ClassificationQutNocturnal@HASOC'19: CNN for Hate Speech and Offensive Content Identification in Hindi Language
We describe our top-team solution to Task 1 for Hindi in the HASOC contest organised by FIRE 2019. The task is to identify hate speech and offensive language in Hindi. More specifically, it is a binary classification pro…
Binary Classification