paper-with-me

홈 › Papers

Stereotypical Bias Removal for Hate Speech Detection Task using Knowledge-based Generalizations

2020-01-15 · Pinkesh Badjatiya, Manish Gupta, Vasudeva Varma

With the ever-increasing cases of hate spread on social media platforms, it is critical to design abuse detection mechanisms to proactively avoid and control such incidents. While there exist methods for hate speech detection, they stereotype words and hence suffer from inherently biased training. Bias removal has been traditionally studied for structured datasets, but we aim at bias mitigation from unstructured text data. In this paper, we make two important contributions. First, we systematically design methods to quantify the bias for any model and propose algorithms for identifying the set of words which the model stereotypes. Second, we propose novel methods leveraging knowledge-based generalizations for bias-free learning. Knowledge-based generalization provides an effective way to encode knowledge because the abstraction they provide not only generalizes content but also facilitates retraction of information from the hate speech detection classifier, thereby reducing the imbalance. We experiment with multiple knowledge generalization policies and analyze their effect on general performance and in mitigating bias. Our experiments with two real-world datasets, a Wikipedia Talk Pages dataset (WikiDetox) of size ~96k and a Twitter dataset of size ~24k, show that the use of knowledge-based generalizations results in better performance by forcing the classifier to learn from generalized content. Our methods utilize existing knowledge-bases and can easily be extended to other tasks

📄 PDF Abstract BibTeX arXiv:2001.05495

Code (0)

등록된 구현이 없습니다.

Tasks

Abuse DetectionHate Speech Detection

Similar Papers 제목 키워드 기반

Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language

2023-10-31 · Jimin Mun, Emily Allaway, Akhila Yerukola, Laura Vianna 외

Counterspeech, i.e., responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship. However, properly countering hateful language …

Philosophy

Improving Counterfactual Generation for Fair Hate Speech Detection

2021-08-03 · ACL (WOAH) 2021 8 · Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy, Mohammad Atari 외

Bias mitigation approaches reduce models' dependence on sensitive features of data, such as social group tokens (SGTs), resulting in equal predictions across the sensitive features. In hate speech detection, however, equ…

counterfactualFairnessHate Speech DetectionSentence

Target Span Detection for Implicit Harmful Content

2024-03-28 · Nazanin Jafari, James Allan, Sheikh Muhammad Sarwar

Identifying the targets of hate speech is a crucial step in grasping the nature of such speech and, ultimately, in improving the detection of offensive posts on online forums. Much harmful content on online platforms use…

Exploring Hate Speech Detection with HateXplain and BERT

2022-08-09 · Arvind Subramaniam, Aryan Mehra, Sayani Kundu

Hate Speech takes many forms to target communities with derogatory comments, and takes humanity a step back in societal progress. HateXplain is a recently published and first dataset to use annotated spans in the form of…

Hate Speech Detection

Darkness can not drive out darkness: Investigating Bias in Hate SpeechDetection Models

2022-05-01 · ACL 2022 5 · Fatma Elsafoury

It has become crucial to develop tools for automated hate speech and abuse detection. These tools would help to stop the bullies and the haters and provide a safer environment for individuals especially from marginalized…

Abuse DetectionHate Speech Detection