paper-with-me

홈 › Papers

Demoting Racial Bias in Hate Speech Detection

2020-05-25 · WS 2020 7 · Mengzhou Xia, Anjalie Field, Yulia Tsvetkov

In current hate speech datasets, there exists a high correlation between annotators' perceptions of toxicity and signals of African American English (AAE). This bias in annotated training data and the tendency of machine learning models to amplify it cause AAE text to often be mislabeled as abusive/offensive/hate speech with a high false positive rate by current hate speech classifiers. In this paper, we use adversarial training to mitigate this bias, introducing a hate speech classifier that learns to detect toxic sentences while demoting confounds corresponding to AAE texts. Experimental results on a hate speech dataset and an AAE dataset suggest that our method is able to substantially reduce the false positive rate for AAE text while only minimally affecting the performance of hate speech classification.

📄 PDF Abstract BibTeX arXiv:2005.12246

Code (0)

등록된 구현이 없습니다.

Tasks

Hate Speech Detection

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

The Risk of Racial Bias in Hate Speech Detection

2019-07-01 · ACL 2019 7 · Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi 외

We investigate how annotators{'} insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations. We first uncover unexp…

Hate Speech Detection

Hate Speech Detection and Racial Bias Mitigation in Social Media based on BERT model

2020-08-14 · Marzieh Mozafari, Reza Farahbakhsh, Noel Crespi

Disparate biases associated with datasets and trained classifiers in hateful and abusive content identification tasks have raised many concerns recently. Although the problem of biased datasets on abusive language detect…

Abusive LanguageHate Speech DetectionLanguage ModellingTransfer Learning

Racial Bias in Hate Speech and Abusive Language Detection Datasets

2019-05-29 · WS 2019 8 · Thomas Davidson, Debasmita Bhattacharya, Ingmar Weber

Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and…

Abuse DetectionAbusive Language

Analyzing Hate Speech Data along Racial, Gender and Intersectional Axes

2022-05-13 · NAACL (GeBNLP) 2022 7 · Antonis Maronikolakis, Philip Baader, Hinrich Schütze

To tackle the rising phenomenon of hate speech, efforts have been made towards data curation and analysis. When it comes to analysis of bias, previous work has focused predominantly on race. In our work, we further inves…

Biasly: a machine learning based platform for automatic racial discrimination detection in online texts

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Detecting hateful, toxic, and otherwise racist or sexist language in user-generated online contents has become an increasingly important task in recent years. Indeed, the anonymity, transience, size of messages, and the …

BIG-bench Machine LearningManagement