paper-with-me

Papers

Detection of Offensive and Threatening Online Content in a Low Resource Language

2023-11-17 · Fatima Muhammad Adam, Abubakar Yakubu Zandam, Isa Inuwa-Dutse

Hausa is a major Chadic language, spoken by over 100 million people in Africa. However, from a computational linguistic perspective, it is considered a low-resource language, with limited resources to support Natural Language Processing (NLP) tasks. Online platforms often facilitate social interactions that can lead to the use of offensive and threatening language, which can go undetected due to the lack of detection systems designed for Hausa. This study aimed to address this issue by (1) conducting two user studies (n=308) to investigate cyberbullying-related issues, (2) collecting and annotating the first set of offensive and threatening datasets to support relevant downstream tasks in Hausa, (3) developing a detection system to flag offensive and threatening content, and (4) evaluating the detection system and the efficacy of the Google-based translation engine in detecting offensive and threatening terms in Hausa. We found that offensive and threatening content is quite common, particularly when discussing religion and politics. Our detection system was able to detect more than 70% of offensive and threatening content, although many of these were mistranslated by Google's translation engine. We attribute this to the subtle relationship between offensive and threatening content and idiomatic expressions in the Hausa language. We recommend that diverse stakeholders participate in understanding local conventions and demographics in order to develop a more effective detection system. These insights are essential for implementing targeted moderation strategies to create a safe and inclusive online environment.

📄 PDF Abstract BibTeX arXiv:2311.10541

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeTranslation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Abusive and Threatening Language Detection in Urdu using Boosting based and BERT based models: A Comparative Approach

2021-11-27 · Mithun Das, Somnath Banerjee, Punyajoy Saha

Online hatred is a growing concern on many social media platforms. To address this issue, different social media platforms have introduced moderation policies for such content. They also employ moderators who can check t…

Abusive Language

Offense Detection in Dravidian Languages using Code-Mixing Index based Focal Loss

2021-11-12 · Debapriya Tula, Shreyas Ms, Viswanatha Reddy, Pranjal Sahu 외

Over the past decade, we have seen exponential growth in online content fueled by social media platforms. Data generation of this scale comes with the caveat of insurmountable offensive content in it. The complexity of i…

Offensive Content Detection via Synthetic Code-Switched Text

2022-10-01 · COLING 2022 10 · Cesa Salaam, Franck Dernoncourt, Trung Bui, Danda Rawat 외

The prevalent use of offensive content in social media has become an important reason for concern for online platforms (customer service chat-boxes, social media platforms, etc). Classifying offensive and hate-speech con…

fr-en

Offensive Content Detection Via Synthetic Code-Switched Text

2022-01-16 · ACL ARR January 2022 1 · Anonymous

The prevalent use of offensive content in social media has become an important reasonfor concern for online platforms (customer service chat-boxes, and social media platforms). Classifying offensive and hate-speech conte…

fr-en

Harnessing Pre-Trained Sentence Transformers for Offensive Language Detection in Indian Languages

2023-10-03 · Ananya Joshi, Raviraj Joshi

In our increasingly interconnected digital world, social media platforms have emerged as powerful channels for the dissemination of hate speech and offensive content. This work delves into the domain of hate speech detec…

Hate Speech DetectionSentencetext-classificationText Classification