paper-with-me

홈 › Papers

Hate Speech and Offensive Language Detection in Bengali

2022-10-07 · Mithun Das, Somnath Banerjee, Punyajoy Saha, Animesh Mukherjee

Social media often serves as a breeding ground for various hateful and offensive content. Identifying such content on social media is crucial due to its impact on the race, gender, or religion in an unprejudiced society. However, while there is extensive research in hate speech detection in English, there is a gap in hateful content detection in low-resource languages like Bengali. Besides, a current trend on social media is the use of Romanized Bengali for regular interactions. To overcome the existing research's limitations, in this study, we develop an annotated dataset of 10K Bengali posts consisting of 5K actual and 5K Romanized Bengali tweets. We implement several baseline models for the classification of such hateful posts. We further explore the interlingual transfer mechanism to boost classification performance. Finally, we perform an in-depth error analysis by looking into the misclassified posts by the models. While training actual and Romanized datasets separately, we observe that XLM-Roberta performs the best. Further, we witness that on joint training and few-shot training, MuRIL outperforms other models by interpreting the semantic expressions better. We make our code and dataset public for others.

📄 PDF Abstract BibTeX arXiv:2210.03479

Code (1)

hate-alert/bengali_hate 공식 구현

Tasks

Hate Speech Detection

Similar Papers 제목 키워드 기반

BanglaHateBERT: BERT for Abusive Language Detection in Bengali

2022-06-01 · ResTUP (LREC) 2022 6 · Md Saroar Jahan, Mainul Haque, Nabil Arhab, Mourad Oussalah

This paper introduces BanglaHateBERT, a retrained BERT model for abusive language detection in Bengali. The model was trained with a large-scale Bengali offensive, abusive, and hateful corpus that we have collected from …

Abusive LanguageLanguage ModelingLanguage Modelling

Harnessing Pre-Trained Sentence Transformers for Offensive Language Detection in Indian Languages

2023-10-03 · Ananya Joshi, Raviraj Joshi

In our increasingly interconnected digital world, social media platforms have emerged as powerful channels for the dissemination of hate speech and offensive content. This work delves into the domain of hate speech detec…

Hate Speech DetectionSentencetext-classificationText Classification

Cross-Linguistic Offensive Language Detection: BERT-Based Analysis of Bengali, Assamese, & Bodo Conversational Hateful Content from Social Media

2023-12-16 · Jhuma Kabir Mim, Mourad Oussalah, Akash Singhal

In today's age, social media reigns as the paramount communication platform, providing individuals with the avenue to express their conjectures, intellectual propositions, and reflections. Unfortunately, this freedom oft…

Language Identification

Hate Speech and Offensive Content Detection in Indo-Aryan Languages: A Battle of LSTM and Transformers

2023-12-09 · Nikhil Narayan, Mrutyunjay Biswal, Pramod Goyal, Abhranta Panigrahi

Social media platforms serve as accessible outlets for individuals to express their thoughts and experiences, resulting in an influx of user-generated data spanning all age groups. While these platforms enable free expre…

Hate Speech DetectionModel SelectionXLM-R

Automated Hate Speech Detection and the Problem of Offensive Language

2017-03-11 · Thomas Davidson, Dana Warmsley, Michael Macy, Ingmar Weber

A key challenge for automatic hate-speech detection on social media is the separation of hate speech from other instances of offensive language. Lexical detection methods tend to have low precision because they classify …

Hate Speech Detection