paper-with-me

Papers

Arabic Offensive Language on Twitter: Analysis and Experiments

2020-04-05 · EACL (WANLP) 2021 4 · Hamdy Mubarak, Ammar Rashed, Kareem Darwish, Younes Samih, Ahmed Abdelali

Detecting offensive language on Twitter has many applications ranging from detecting/predicting bullying to measuring polarization. In this paper, we focus on building a large Arabic offensive tweet dataset. We introduce a method for building a dataset that is not biased by topic, dialect, or target. We produce the largest Arabic dataset to date with special tags for vulgarity and hate speech. We thoroughly analyze the dataset to determine which topics, dialects, and gender are most associated with offensive tweets and how Arabic speakers use offensive language. Lastly, we conduct many experiments to produce strong results (F1 = 83.2) on the dataset using SOTA techniques.

📄 PDF Abstract BibTeX arXiv:2004.02192

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emojis as Anchors to Detect Arabic Offensive Language and Hate Speech

2022-01-18 · Hamdy Mubarak, Sabit Hassan, Shammur Absar Chowdhury

We introduce a generic, language-independent method to collect a large percentage of offensive and hate tweets regardless of their topics or genres. We harness the extralinguistic information embedded in the emojis to co…

Cultural Vocal Bursts Intensity Prediction

CyberTronics at SemEval-2020 Task 12: Multilingual Offensive Language Identification over Social Media

2020-12-01 · SEMEVAL 2020 · Sayanta Paul, Sriparna Saha, Mohammed Hasanuzzaman

The SemEval-2020 Task 12 (OffensEval) challenge focuses on detection of signs of offensiveness using posts or comments over social media. This task has been organized for several languages, e.g., Arabic, Danish, English,…

Feature EngineeringLanguage Identification

LISAC FSDM-USMBA Team at SemEval-2020 Task 12: Overcoming AraBERT's pretrain-finetune discrepancy for Arabic offensive language identification

2020-12-01 · SEMEVAL 2020 · Hamza Alami, Said Ouatik El Alaoui, Abdessamad Benlahbib, Noureddine En-nahnahi

AraBERT is an Arabic version of the state-of-the-art Bidirectional Encoder Representations from Transformers (BERT) model. The latter has achieved good performance in a variety of Natural Language Processing (NLP) tasks.…

Language Identification

Automatic Expansion and Retargeting of Arabic Offensive Language Training

2021-11-18 · Hamdy Mubarak, Ahmed Abdelali, Kareem Darwish, Younes Samih

Rampant use of offensive language on social media led to recent efforts on automatic identification of such language. Though offensive language has general characteristics, attacks on specific entities may exhibit distin…

Abusive Language Detection on Arabic Social Media

2017-08-01 · WS 2017 8 · Hamdy Mubarak, Kareem Darwish, Walid Magdy

In this paper, we present our work on detecting abusive language on Arabic social media. We extract a list of obscene words and hashtags using common patterns used in offensive and rude communications. We also classify T…

Abusive LanguageGeneral Classification