Mitigating Bias in Conversations: A Hate Speech Classifier and Debiaser with Prompts
Discriminatory language and biases are often present in hate speech during conversations, which usually lead to negative impacts on targeted groups such as those based on race, gender, and religion. To tackle this issue, we propose an approach that involves a two-step process: first, detecting hate speech using a classifier, and then utilizing a debiasing component that generates less biased or unbiased alternatives through prompts. We evaluated our approach on a benchmark dataset and observed reduction in negativity due to hate speech comments. The proposed method contributes to the ongoing efforts to reduce biases in online discourse and promote a more inclusive and fair environment for communication.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Countering hate on social media: Large scale classification of hate and counter speech
Hateful rhetoric is plaguing online discourse, fostering extreme societal movements and possibly giving rise to real-world violence. A potential solution to this growing global problem is citizen-generated counter speech…
Ensemble LearningGeneral ClassificationImpact of Politically Biased Data on Hate Speech Classification
One challenge that social media platforms are facing nowadays is hate speech. Hence, automatic hate speech detection has been increasingly researched in recent years - in particular with the rise of deep learning. A prob…
ClassificationHate Speech DetectionStereotypical Bias Removal for Hate Speech Detection Task using Knowledge-based Generalizations
With the ever-increasing cases of hate spread on social media platforms, it is critical to design abuse detection mechanisms to proactively avoid and control such incidents. While there exist methods for hate speech dete…
Abuse DetectionHate Speech DetectionDemoting Racial Bias in Hate Speech Detection
In current hate speech datasets, there exists a high correlation between annotators' perceptions of toxicity and signals of African American English (AAE). This bias in annotated training data and the tendency of machine…
Hate Speech DetectionPower of Explanations: Towards automatic debiasing in hate speech detection
Hate speech detection is a common downstream application of natural language processing (NLP) in the real world. In spite of the increasing accuracy, current data-driven approaches could easily learn biases from the imba…
FairnessHate Speech Detection