Towards Automatic Online Hate Speech Intervention Generation using Pretrained Language Model
Social media harbours substantial toxic and hateful conversations today. Curbing them has emerged as a critical challenge for governments and organizations globally. Prior research has primarily concentrated on the detection of online hate speech while ignoring further action needed to discourage individuals from using hate speech in the future. Counterspeech is an effective way to tackle online hate, leaving freedom of speech untouched. The focus is to directly intervene in the conversation with textual responses that counter the hate content and prevent it from further spreading. In this paper, we propose a novel natural language generation task for hate speech intervention, where the goal is to automatically generate responses to intervene during online conversations that contain hate speech. We sequentially analyzed the performance and capability of various state-of-the-art pretrained language models dialogue generation model for automated hate speech intervention system using automatic metric and manual human evaluation. The results indicate that the generated intervention responses are very promising in terms of relevance and contextual meaning
Code (0)
등록된 구현이 없습니다.
Tasks
Dialogue GenerationLanguage ModelingLanguage ModellingText GenerationSimilar Papers 제목 키워드 기반
A Benchmark Dataset for Learning to Intervene in Online Hate Speech
Countering online hate speech is a critical yet challenging task, but one which can be aided by the use of Natural Language Processing (NLP) techniques. Previous research has primarily focused on the development of NLP m…
Response GenerationTowards Automatic Generation of Messages Countering Online Hate Speech and Microaggressions
With the widespread use of social media, online hate is increasing, and microaggressions are receiving attention. We explore the potential for using pretrained language models to automatically generate messages that comb…
InformativenessHuman-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech Countering
Fighting online hate speech is a challenge that is usually addressed using Natural Language Processing via automatic detection and removal of hate content. Besides this approach, counter narratives have emerged as an eff…
Text GenerationCountering Online Hate Speech: An NLP Perspective
Online hate speech has caught everyone's attention from the news related to the COVID-19 pandemic, US elections, and worldwide protests. Online toxicity - an umbrella term for online hateful behavior, manifests itself in…
Towards Knowledge-Grounded Counter Narrative Generation for Hate Speech
Tackling online hatred using informed textual responses - called counter narratives - has been brought under the spotlight recently. Accordingly, a research line has emerged to automatically generate counter narratives i…