Examining Racial Bias in an Online Abuse Corpus with Structural Topic Modeling
We use structural topic modeling to examine racial bias in data collected to train models to detect hate speech and abusive language in social media posts. We augment the abusive language dataset by adding an additional feature indicating the predicted probability of the tweet being written in African-American English. We then use structural topic modeling to examine the content of the tweets and how the prevalence of different topics is related to both abusiveness annotation and dialect prediction. We find that certain topics are disproportionately racialized and considered abusive. We discuss how topic modeling may be a useful approach for identifying bias in annotated data.
Code (1)
Tasks
Abusive LanguageSimilar Papers 제목 키워드 기반
Towards a Comprehensive Taxonomy and Large-Scale Annotated Corpus for Online Slur Usage
Abusive language classifiers have been shown to exhibit bias against women and racial minorities. Since these models are trained on data that is collected using keywords, they tend to exhibit a high sensitivity towards p…
Abusive LanguageInvestigating Sports Commentator Bias within a Large Corpus of American Football Broadcasts
Sports broadcasters inject drama into play-by-play commentary by building team and player narratives through subjective analyses and anecdotes. Prior studies based on small datasets and manual coding show that such theat…
Racial Bias in Hate Speech and Abusive Language Detection Datasets
Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and…
Abuse DetectionAbusive LanguageA Longitudinal Analysis of Racial and Gender Bias in New York Times and Fox News Images and Articles
The manner in which different racial and gender groups are portrayed in news coverage plays a large role in shaping public opinion. As such, understanding how such groups are portrayed in news media is of notable societa…
ArticlesStudying Bias in GANs through the Lens of Race
In this work, we study how the performance and evaluation of generative image models are impacted by the racial composition of their training datasets. By examining and controlling the racial distributions in various tra…