Investigating Annotator Bias in Abusive Language Datasets
Nowadays, social media platforms use classification models to cope with hate speech and abusive language. The problem of these models is their vulnerability to bias. A prevalent form of bias in hate speech and abusive language datasets is annotator bias caused by the annotator’s subjective perception and the complexity of the annotation task. In our paper, we develop a set of methods to measure annotator bias in abusive language datasets and to identify different perspectives on abusive language. We apply these methods to four different abusive language datasets. Our proposed approach supports annotation processes of such datasets and future research addressing different perspectives on the perception of abusive language.
Code (1)
Tasks
Abusive LanguageSimilar Papers 제목 키워드 기반
Investigating Sampling Bias in Abusive Language Detection
Abusive language detection is becoming increasingly important, but we still understand little about the biases in our datasets for abusive language detection, and how these biases affect the quality of abusive language d…
Abusive LanguageLower Bias, Higher Density Abusive Language Datasets: A Recipe
Datasets to train models for abusive language detection are at the same time necessary and still scarce. One the reasons for their limited availability is the cost of their creation. It is not only that manual annotation…
Abusive LanguageIdentifying and Measuring Annotator Bias Based on Annotators’ Demographic Characteristics
Machine learning is recently used to detect hate speech and other forms of abusive language in online platforms. However, a notable weakness of machine learning models is their vulnerability to bias, which can impair the…
Abusive LanguageBIG-bench Machine LearningFairnessRacial Bias in Hate Speech and Abusive Language Detection Datasets
Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and…
Abuse DetectionAbusive LanguageToward Annotator Group Bias in Crowdsourcing
Crowdsourcing has emerged as a popular approach for collecting annotated data to train supervised machine learning models. However, annotator bias can lead to defective annotations. Though there are a few works investiga…