An Annotated Corpus of Arabic Tweets for Hate Speech Analysis
Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and annotated each tweet, whether it contains offensive content or not. If a text contains offensive content, we further classify it into different hate speech targets such as religion, gender, politics, ethnicity, origin, and others. A text can contain either single or multiple targets. Multiple annotators are involved in the data annotation task. We calculated the inter-annotator agreement, which was reported to be 0.86 for offensive content and 0.71 for multiple hate speech targets. Finally, we evaluated the data annotation task by employing a different transformers-based model in which AraBERTv2 outperformed with a micro-F1 score of 0.7865 and an accuracy of 0.786.
Code (1)
Similar Papers 제목 키워드 기반
Hate Speech Detection in Saudi Twittersphere: A Deep Learning Approach
With the rise of hate speech phenomena in Twittersphere, significant research efforts have been undertaken to provide automatic solutions for detecting hate speech, varying from simple ma-chine learning models to more co…
Deep LearningHate Speech DetectionAra-Women-Hate: An Annotated Corpus Dedicated to Hate Speech Detection against Women in the Arabic Community
In this paper, an approach for hate speech detection against women in the Arabic community on social media (e.g. Youtube) is proposed. In the literature, similar works have been presented for other languages such as Engl…
Hate Speech DetectionEnsemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
Today, hate speech classification from Arabic tweets has drawn the attention of several researchers. Many systems and techniques have been developed to resolve this classification task. Nevertheless, two of the major cha…
Data AugmentationEnsemble LearningHate Speech DetectionExamining a hate speech corpus for hate speech detection and popularity prediction
As research on hate speech becomes more and more relevant every day, most of it is still focused on hate speech detection. By attempting to replicate a hate speech detection experiment performed on an existing Twitter co…
Hate Speech DetectionAraCOVID19-MFH: Arabic COVID-19 Multi-label Fake News and Hate Speech Detection Dataset
Along with the COVID-19 pandemic, an "infodemic" of false and misleading information has emerged and has complicated the COVID-19 response efforts. Social networking sites such as Facebook and Twitter have contributed la…
ArticlesDialect IdentificationFact CheckingFake News Detection+2