Intersectional Bias in Hate Speech and Abusive Language Datasets
Algorithms are widely applied to detect hate speech and abusive language in social media. We investigated whether the human-annotated data used to train these algorithms are biased. We utilized a publicly available annotated Twitter dataset (Founta et al. 2018) and classified the racial, gender, and party identification dimensions of 99,996 tweets. The results showed that African American tweets were up to 3.7 times more likely to be labeled as abusive, and African American male tweets were up to 77% more likely to be labeled as hateful compared to the others. These patterns were statistically significant and robust even when party identification was added as a control variable. This study provides the first systematic evidence on intersectional bias in datasets of hate speech and abusive language.
Code (1)
Tasks
Abuse DetectionAbusive LanguageMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Investigating Annotator Bias in Abusive Language Datasets
Nowadays, social media platforms use classification models to cope with hate speech and abusive language. The problem of these models is their vulnerability to bias. A prevalent form of bias in hate speech and abusive la…
Abusive LanguageRacial Bias in Hate Speech and Abusive Language Detection Datasets
Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and…
Abuse DetectionAbusive LanguageAnalyzing Hate Speech Data along Racial, Gender and Intersectional Axes
To tackle the rising phenomenon of hate speech, efforts have been made towards data curation and analysis. When it comes to analysis of bias, previous work has focused predominantly on race. In our work, we further inves…
Hate Speech Detection and Racial Bias Mitigation in Social Media based on BERT model
Disparate biases associated with datasets and trained classifiers in hateful and abusive content identification tasks have raised many concerns recently. Although the problem of biased datasets on abusive language detect…
Abusive LanguageHate Speech DetectionLanguage ModellingTransfer LearningL-HSAB: A Levantine Twitter Dataset for Hate Speech and Abusive Language
Hate speech and abusive language have become a common phenomenon on Arabic social media. Automatic hate speech and abusive detection systems can facilitate the prohibition of toxic textual contents. The complexity, infor…
Abusive LanguageHate Speech Detection