Demoting Racial Bias in Hate Speech Detection
In current hate speech datasets, there exists a high correlation between annotators' perceptions of toxicity and signals of African American English (AAE). This bias in annotated training data and the tendency of machine learning models to amplify it cause AAE text to often be mislabeled as abusive/offensive/hate speech with a high false positive rate by current hate speech classifiers. In this paper, we use adversarial training to mitigate this bias, introducing a hate speech classifier that learns to detect toxic sentences while demoting confounds corresponding to AAE texts. Experimental results on a hate speech dataset and an AAE dataset suggest that our method is able to substantially reduce the false positive rate for AAE text while only minimally affecting the performance of hate speech classification.
Code (0)
등록된 구현이 없습니다.
Tasks
Hate Speech DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
The Risk of Racial Bias in Hate Speech Detection
We investigate how annotators{'} insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations. We first uncover unexp…
Hate Speech DetectionHate Speech Detection and Racial Bias Mitigation in Social Media based on BERT model
Disparate biases associated with datasets and trained classifiers in hateful and abusive content identification tasks have raised many concerns recently. Although the problem of biased datasets on abusive language detect…
Abusive LanguageHate Speech DetectionLanguage ModellingTransfer LearningRacial Bias in Hate Speech and Abusive Language Detection Datasets
Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and…
Abuse DetectionAbusive LanguageAnalyzing Hate Speech Data along Racial, Gender and Intersectional Axes
To tackle the rising phenomenon of hate speech, efforts have been made towards data curation and analysis. When it comes to analysis of bias, previous work has focused predominantly on race. In our work, we further inves…
Biasly: a machine learning based platform for automatic racial discrimination detection in online texts
Detecting hateful, toxic, and otherwise racist or sexist language in user-generated online contents has become an increasingly important task in recent years. Indeed, the anonymity, transience, size of messages, and the …
BIG-bench Machine LearningManagement