paper-with-me

Papers

Outcome-Constrained Large Language Models for Countering Hate Speech

2024-03-25 · Lingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying Song

Automatic counterspeech generation methods have been developed to assist efforts in combating hate speech. Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven. However, the real impact of counterspeech in online environments is seldom considered. This study aims to develop methods for generating counterspeech constrained by conversation outcomes and evaluate their effectiveness. We experiment with large language models (LLMs) to incorporate into the text generation process two desired conversation outcomes: low conversation incivility and non-hateful hater reentry. Specifically, we experiment with instruction prompts, LLM finetuning, and LLM reinforcement learning (RL). Evaluation results show that our methods effectively steer the generation of counterspeech toward the desired outcomes. Our analyses, however, show that there are differences in the quality and style depending on the model.

📄 PDF Abstract BibTeX arXiv:2403.17146

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)Text Generation

Similar Papers 제목 키워드 기반

Towards countering hate speech against journalists on social media

2019-12-05 · Polychronis Charitidis, Stavros Doropoulos, Stavros Vologiannidis, Ioannis Papastergiou 외

The damaging effects of hate speech on social media are evident during the last few years, and several organizations, researchers and social media platforms tried to harness them in various ways. Despite these efforts, s…

Active LearningHate Speech Detection

Countering Online Hate Speech: An NLP Perspective

2021-09-07 · Mudit Chaudhary, Chandni Saxena, Helen Meng

Online hate speech has caught everyone's attention from the news related to the COVID-19 pandemic, US elections, and worldwide protests. Online toxicity - an umbrella term for online hateful behavior, manifests itself in…

Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language

2023-10-31 · Jimin Mun, Emily Allaway, Akhila Yerukola, Laura Vianna 외

Counterspeech, i.e., responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship. However, properly countering hateful language …

Philosophy

Towards Automatic Generation of Messages Countering Online Hate Speech and Microaggressions

2022-07-01 · NAACL (WOAH) 2022 7 · Mana Ashida, Mamoru Komachi

With the widespread use of social media, online hate is increasing, and microaggressions are receiving attention. We explore the potential for using pretrained language models to automatically generate messages that comb…

Informativeness

Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues

2024-01-15 · Sougata Saha, Rohini Srihari

Hateful comments are prevalent on social media platforms. Although tools for automatically detecting, flagging, and blocking such false, offensive, and harmful content online have lately matured, such reactive and brute …

BlockingResponse Generation