paper-with-me

홈 › Papers

A Large-scale Comprehensive Abusiveness Detection Dataset with Multifaceted Labels from Reddit

2021-11-01 · CoNLL (EMNLP) 2021 11 · Hoyun Song, Soo Hyun Ryu, Huije Lee, Jong Park

As users in online communities suffer from severe side effects of abusive language, many researchers attempted to detect abusive texts from social media, presenting several datasets for such detection. However, none of them contain both comprehensive labels and contextual information, which are essential for thoroughly detecting all kinds of abusiveness from texts, since datasets with such fine-grained features demand a significant amount of annotations, leading to much increased complexity. In this paper, we propose a Comprehensive Abusiveness Detection Dataset (CADD), collected from the English Reddit posts, with multifaceted labels and contexts. Our dataset is annotated hierarchically for an efficient annotation through crowdsourcing on a large-scale. We also empirically explore the characteristics of our dataset and provide a detailed analysis for novel insights. The results of our experiments with strong pre-trained natural language understanding models on our dataset show that our dataset gives rise to meaningful performance, assuring its practicality for abusive language detection.

📄 PDF Abstract BibTeX

Code (1)

nlpcl-lab/cadd_dataset 공식 구현

Tasks

Abusive LanguageNatural Language Understanding

Similar Papers 제목 키워드 기반

A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings

2025-05-17 · Fitsum Gaim, Hoyun Song, Huije Lee, Changgeon Ko 외

Content moderation research has recently made significant advances, but still fails to serve the majority of the world's languages due to the lack of resources, leaving millions of vulnerable users to online hostility. T…

Abusive LanguageTopic Classification

Detecting context abusiveness using hierarchical deep learning

2019-11-01 · WS 2019 11 · Ju-Hyoung Lee, Jun-U Park, Jeong-Won Cha, Yo-Sub Han

Abusive text is a serious problem in social media and causes many issues among users as the number of users and the content volume increase. There are several attempts for detecting or preventing abusive text effectively…

Deep Learning

Unraveling the Search Space of Abusive Language in Wikipedia with Dynamic Lexicon Acquisition

2019-11-01 · WS 2019 11 · Wei-Fan Chen, Khalid Al Khatib, Matthias Hagen, Henning Wachsmuth 외

Many discussions on online platforms suffer from users offending others by using abusive terminology, threatening each other, or being sarcastic. Since an automatic detection of abusive language can support human moderat…

Abusive Language

Multilingual Abusiveness Identification on Code-Mixed Social Media Text

2022-03-01 · Ekagra Ranjan, Naman Poddar

Social Media platforms have been seeing adoption and growth in their usage over time. This growth has been further accelerated with the lockdown in the past year when people's interaction, conversation, and expression we…

SentenceTransliteration

Do You Really Want to Hurt Me? Predicting Abusive Swearing in Social Media

2020-05-01 · LREC 2020 5 · Endang Wahyu Pamungkas, Valerio Basile, Viviana Patti

Swearing plays an ubiquitous role in everyday conversations among humans, both in oral and textual communication, and occurs frequently in social media texts, typically featured by informal language and spontaneous writi…