paper-with-me

홈 › Papers

BAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl web service

2023-08-21 · Anna Kołos, Inez Okulska, Kinga Głąbińska, Agnieszka Karlińska, Emilia Wiśnios, Paweł Ellerik, Andrzej Prałat

Since the Internet is flooded with hate, it is one of the main tasks for NLP experts to master automated online content moderation. However, advancements in this field require improved access to publicly available accurate and non-synthetic datasets of social media content. For the Polish language, such resources are very limited. In this paper, we address this gap by presenting a new open dataset of offensive social media content for the Polish language. The dataset comprises content from Wykop.pl, a popular online service often referred to as the "Polish Reddit", reported by users and banned in the internal moderation process. It contains a total of 691,662 posts and comments, evenly divided into two categories: "harmful" and "neutral" ("non-harmful"). The anonymized subset of the BAN-PL dataset consisting on 24,000 pieces (12,000 for each class), along with preprocessing scripts have been made publicly available. Furthermore the paper offers valuable insights into real-life content moderation processes and delves into an analysis of linguistic features and content characteristics of the dataset. Moreover, a comprehensive anonymization procedure has been meticulously described and applied. The prevalent biases encountered in similar datasets, including post-moderation and pre-selection biases, are also discussed.

📄 PDF Abstract BibTeX arXiv:2308.10592

Code (1)

ziliat-nask/ban-pl 공식 구현

Tasks

Specificity

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

HateBERT: Retraining BERT for Abusive Language Detection in English

2020-10-23 · ACL (WOAH) 2021 8 · Tommaso Caselli, Valerio Basile, Jelena Mitrović, Michael Granitzer

In this paper, we introduce HateBERT, a re-trained BERT model for abusive language detection in English. The model was trained on RAL-E, a large-scale dataset of Reddit comments in English from communities banned for bei…

Abusive LanguageHate Speech DetectionLanguage ModelingLanguage Modelling

StopHC: A Harmful Content Detection and Mitigation Architecture for Social Media Platforms

2024-11-09 · Ciprian-Octavian Truică, Ana-Teodora Constantinescu, Elena-Simona Apostol

The mental health of social media users has started more and more to be put at risk by harmful, hateful, and offensive content. In this paper, we propose \textsc{StopHC}, a harmful content detection and mitigation archit…

The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation

2025-05-31 · YuHang Zhou, Yimin Xiao, Wei Ai, Ge Gao

Social media platforms have become central to modern communication, yet they also harbor offensive content that challenges platform safety and inclusivity. While prior research has primarily focused on textual indicators…

"HOT" ChatGPT: The promise of ChatGPT in detecting and discriminating hateful, offensive, and toxic comments on social media

2023-04-20 · Lingyao Li, Lizhou Fan, Shubham Atreja, Libby Hemphill

Harmful content is pervasive on social media, poisoning online communities and negatively impacting participation. A common approach to address this issue is to develop detection models that rely on human annotations. Ho…

HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models

2024-03-18 · Huy Nghiem, Hal Daumé III

The widespread use of social media necessitates reliable and efficient detection of offensive content to mitigate harmful effects. Although sophisticated models perform well on individual datasets, they often fail to gen…