Directions in Abusive Language Training Data: Garbage In, Garbage Out
Data-driven analysis and detection of abusive online content covers many different tasks, phenomena, contexts, and methodologies. This paper systematically reviews abusive language dataset creation and content in conjunction with an open website for cataloguing abusive language data. This collection of knowledge leads to a synthesis providing evidence-based recommendations for practitioners working with this complex and highly diverse data.
Code (0)
등록된 구현이 없습니다.
Tasks
Abusive LanguageSimilar Papers 제목 키워드 기반
"To Target or Not to Target": Identification and Analysis of Abusive Text Using Ensemble of Classifiers
With rising concern around abusive and hateful behavior on social media platforms, we present an ensemble learning method to identify and analyze the linguistic properties of such content. Our stacked ensemble comprises …
BIG-bench Machine LearningEnsemble LearningGarbage in, garbage out: Zero-shot detection of crime using Large Language Models
This paper proposes exploiting the common sense knowledge learned by large language models to perform zero-shot reasoning about crimes given textual descriptions of surveillance videos. We show that when video is (manual…
Common Sense ReasoningLanguage ModelingLanguage ModellingLarge Language ModelCross-domain and Cross-lingual Abusive Language Detection: A Hybrid Approach with Deep Learning and a Multilingual Lexicon
The development of computational methods to detect abusive language in social media within variable and multilingual contexts has recently gained significant traction. The growing interest is confirmed by the large numbe…
Abusive LanguageHateBERT: Retraining BERT for Abusive Language Detection in English
In this paper, we introduce HateBERT, a re-trained BERT model for abusive language detection in English. The model was trained on RAL-E, a large-scale dataset of Reddit comments in English from communities banned for bei…
Abusive LanguageHate Speech DetectionLanguage ModelingLanguage ModellingMulti-Stage Training for Abusive Comment Detection in Indic Languages
In recent years social media has become an increasingly popular tool for communication. People use it to share their ideas, exchange information, and discuss thoughts. Given its prevalence and widespread reach, social me…