paper-with-me

홈 › Papers

Lower Bias, Higher Density Abusive Language Datasets: A Recipe

2020-05-01 · LREC 2020 5 · Juliet van Rosendaal, Tommaso Caselli, Malvina Nissim

Datasets to train models for abusive language detection are at the same time necessary and still scarce. One the reasons for their limited availability is the cost of their creation. It is not only that manual annotation is expensive, it is also the case that the phenomenon is sparse, causing human annotators having to go through a large number of irrelevant examples in order to obtain some significant data. Strategies used until now to increase density of abusive language and obtain more meaningful data overall, include data filtering on the basis of pre-selected keywords and hate-rich sources of data. We suggest a recipe that at the same time can provide meaningful data with possibly higher density of abusive language and also reduce top-down biases imposed by corpus creators in the selection of the data to annotate. More specifically, we exploit the controversy channel on Reddit to obtain keywords that are used to filter a Twitter dataset. While the method needs further validation and refinement, our preliminary experiments show a higher density of abusive tweets in the filtered vs unfiltered dataset, and a more meaningful topic distribution after filtering.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Abusive Language

Similar Papers 제목 키워드 기반

Detection of Abusive Language: the Problem of Biased Datasets

2019-06-01 · NAACL 2019 6

We discuss the impact of data bias on abusive language detection. We show that classification scores on popular datasets reported in previous work are much lower under realistic settings in which this bias is reduced. Su…

Abusive Language

Racial Bias in Hate Speech and Abusive Language Detection Datasets

2019-05-29 · WS 2019 8 · Thomas Davidson, Debasmita Bhattacharya, Ingmar Weber

Technologies for abusive language detection are being developed and applied with little consideration of their potential biases. We examine racial bias in five different sets of Twitter data annotated for hate speech and…

Abuse DetectionAbusive Language

Investigating Annotator Bias in Abusive Language Datasets

2021-09-01 · RANLP 2021 9 · Maximilian Wich, Christian Widmer, Gerhard Hagerer, Georg Groh

Nowadays, social media platforms use classification models to cope with hate speech and abusive language. The problem of these models is their vulnerability to bias. A prevalent form of bias in hate speech and abusive la…

Abusive Language

Examining Temporal Bias in Abusive Language Detection

2023-09-25 · Mali Jin, Yida Mu, Diana Maynard, Kalina Bontcheva

The use of abusive language online has become an increasingly pervasive problem that damages both individuals and society, with effects ranging from psychological harm right through to escalation to real-life violence an…

Abusive Language

Investigating Sampling Bias in Abusive Language Detection

2020-11-01 · EMNLP (ALW) 2020 11 · Dante Razo, Sandra Kübler

Abusive language detection is becoming increasingly important, but we still understand little about the biases in our datasets for abusive language detection, and how these biases affect the quality of abusive language d…

Abusive Language