Cross-Domain Detection of Abusive Language Online
We investigate to what extent the models trained to detect general abusive language generalize between different datasets labeled with different abusive language types. To this end, we compare the cross-domain performance of simple classification models on nine different datasets, finding that the models fail to generalize to out-domain datasets and that having at least some in-domain data is important. We also show that using the frustratingly simple domain adaptation (Daume III, 2007) in most cases improves the results over in-domain training, especially when used to augment a smaller dataset with a larger one.
Code (0)
등록된 구현이 없습니다.
Tasks
Abusive LanguageDomain AdaptationGeneral ClassificationSimilar Papers 제목 키워드 기반
Detect All Abuse! Toward Universal Abusive Language Detection Models
Online abusive language detection (ALD) has become a societal issue of increasing importance in recent years. Several previous works in online ALD focused on solving a single abusive language problem in a single domain, …
Abusive LanguageAllGraph EmbeddingCross-Platform and Cross-Domain Abusive Language Detection with Supervised Contrastive Learning
The prevalence of abusive language on different online platforms has been a major concern that raises the need for automated cross-platform abusive language detection. However, prior works focus on concatenating data fro…
Abusive LanguageContrastive LearningDomain GeneralizationMeta-LearningXHate-999: Analyzing and Detecting Abusive Language Across Domains and Languages
We present XHate-999, a multi-domain and multilingual evaluation data set for abusive language detection. By aligning test instances across six typologically diverse languages, XHate-999 for the first time allows for dis…
Abusive LanguageDisentanglementLanguage ModelingLanguage Modelling+1Cross-domain and Cross-lingual Abusive Language Detection: A Hybrid Approach with Deep Learning and a Multilingual Lexicon
The development of computational methods to detect abusive language in social media within variable and multilingual contexts has recently gained significant traction. The growing interest is confirmed by the large numbe…
Abusive LanguageAbusive and Threatening Language Detection in Urdu using Boosting based and BERT based models: A Comparative Approach
Online hatred is a growing concern on many social media platforms. To address this issue, different social media platforms have introduced moderation policies for such content. They also employ moderators who can check t…
Abusive Language