Outcome-Constrained Large Language Models for Countering Hate Speech
Automatic counterspeech generation methods have been developed to assist efforts in combating hate speech. Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven. However, the real impact of counterspeech in online environments is seldom considered. This study aims to develop methods for generating counterspeech constrained by conversation outcomes and evaluate their effectiveness. We experiment with large language models (LLMs) to incorporate into the text generation process two desired conversation outcomes: low conversation incivility and non-hateful hater reentry. Specifically, we experiment with instruction prompts, LLM finetuning, and LLM reinforcement learning (RL). Evaluation results show that our methods effectively steer the generation of counterspeech toward the desired outcomes. Our analyses, however, show that there are differences in the quality and style depending on the model.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement Learning (RL)Text GenerationSimilar Papers 제목 키워드 기반
Towards countering hate speech against journalists on social media
The damaging effects of hate speech on social media are evident during the last few years, and several organizations, researchers and social media platforms tried to harness them in various ways. Despite these efforts, s…
Active LearningHate Speech DetectionCountering Online Hate Speech: An NLP Perspective
Online hate speech has caught everyone's attention from the news related to the COVID-19 pandemic, US elections, and worldwide protests. Online toxicity - an umbrella term for online hateful behavior, manifests itself in…
Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language
Counterspeech, i.e., responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship. However, properly countering hateful language …
PhilosophyTowards Automatic Generation of Messages Countering Online Hate Speech and Microaggressions
With the widespread use of social media, online hate is increasing, and microaggressions are receiving attention. We explore the potential for using pretrained language models to automatically generate messages that comb…
InformativenessConsolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
Hateful comments are prevalent on social media platforms. Although tools for automatically detecting, flagging, and blocking such false, offensive, and harmful content online have lately matured, such reactive and brute …
BlockingResponse Generation