paper-with-me

홈 › Papers

Bangla Text Dataset and Exploratory Analysis for Online Harassment Detection

2021-02-04 · Md Faisal Ahmed, Zalish Mahmud, Zarin Tasnim Biash, Ahmed Ann Noor Ryen, Arman Hossain, Faisal Bin Ashraf

Being the seventh most spoken language in the world, the use of the Bangla language online has increased in recent times. Hence, it has become very important to analyze Bangla text data to maintain a safe and harassment-free online place. The data that has been made accessible in this article has been gathered and marked from the comments of people in public posts by celebrities, government officials, athletes on Facebook. The total amount of collected comments is 44001. The dataset is compiled with the aim of developing the ability of machines to differentiate whether a comment is a bully expression or not with the help of Natural Language Processing and to what extent it is improper if it is an inappropriate comment. The comments are labeled with different categories of harassment. Exploratory analysis from different perspectives is also included in this paper to have a detailed overview. Due to the scarcity of data collection of categorized Bengali language comments, this dataset can have a significant role for research in detecting bully words, identifying inappropriate comments, detecting different categories of Bengali bullies, etc. The dataset is publicly available at https://data.mendeley.com/datasets/9xjx8twk8p.

📄 PDF Abstract BibTeX arXiv:2102.02478

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BanglaCoNER: Towards Robust Bangla Complex Named Entity Recognition

2023-03-16 · HAZ Sameen Shahgir, Ramisa Alam, Md. Zarif Ul Alam

Named Entity Recognition (NER) is a fundamental task in natural language processing that involves identifying and classifying named entities in text. But much work hasn't been done for complex named entity recognition in…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing

2024-10-22 · Azmine Toushik Wasi, Wahid Faisal, Taj Ahmad, Abdur Rahman 외

Climate change poses critical challenges globally, disproportionately affecting low-income countries that often lack resources and linguistic representation on the international stage. Despite Bangladesh's status as one …

ArticlesAuthority Involvement DetectionData Sources DetectionDetecting Climate/Environmental Topics+7

Sentiment Analysis on Bangla and Romanized Bangla Text (BRBT) using Deep Recurrent models

2016-10-02 · A. Hassan, M. R. Amin, N. Mohammed, A. K. A. Azad

Sentiment Analysis (SA) is an action research area in the digital age. With rapid and constant growth of online social media sites and services, and the increasing amount of textual data such as - statuses, comments, rev…

Sentiment Analysis

Mukh-Oboyob: Stable Diffusion and BanglaBERT enhanced Bangla Text-to-Face Synthesis

2023-11-01 · International Journal of Advanced Computer Science and Applications 2023 11 · Aloke Kumar Saha, Noor Mairukh Khan Arnob, Nakiba Nuren Rahman, Maria Haque 외

Facial image generation from textual generation is one of the most complicated tasks within the broader topic of Text-to-Image (TTI) synthesis. It is relevant in several fields of scientific research, cartoon and animati…

Face GenerationImage GenerationLanguage ModelingLanguage Modelling+2

Crime Prediction using Machine Learning with a Novel Crime Dataset

2022-11-03 · Faisal Tareque Shohan, Abu Ubaida Akash, Muhammad Ibrahim, Mohammad Shafiul Alam

Crime is an unlawful act that carries legal repercussions. Bangladesh has a high crime rate due to poverty, population growth, and many other socio-economic issues. For law enforcement agencies, understanding crime patte…

ArticlesCrime Prediction