paper-with-me

Papers

Towards Automatic Generation of Messages Countering Online Hate Speech and Microaggressions

2022-07-01 · NAACL (WOAH) 2022 7 · Mana Ashida, Mamoru Komachi

With the widespread use of social media, online hate is increasing, and microaggressions are receiving attention. We explore the potential for using pretrained language models to automatically generate messages that combat the associated offensive texts. Specifically, we focus on using prompting to steer model generation as it requires less data and computation than fine-tuning. We also propose a human evaluation perspective; offensiveness, stance, and informativeness. After obtaining 306 counterspeech and 42 microintervention messages generated by GPT-{2, 3, Neo}, we conducted a human evaluation using Amazon Mechanical Turk. The results indicate the potential of using prompting in the proposed generation task. All the generated texts along with the annotation are published to encourage future research on countering hate and microaggressions online.

📄 PDF Abstract BibTeX

Code (1)

tmu-nlp/chasm 공식 구현

Tasks

Informativeness

Similar Papers 제목 키워드 기반

Empowering NGOs in Countering Online Hate Messages

2021-07-06 · Yi-Ling Chung, Serra Sinem Tekiroglu, Sara Tonelli, Marco Guerini

Studies on online hate speech have mostly focused on the automated detection of harmful messages. Little attention has been devoted so far to the development of effective strategies to fight hate speech, in particular th…

Management

A Benchmark Dataset for Learning to Intervene in Online Hate Speech

2019-09-10 · IJCNLP 2019 11 · Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding 외

Countering online hate speech is a critical yet challenging task, but one which can be aided by the use of Natural Language Processing (NLP) techniques. Previous research has primarily focused on the development of NLP m…

Response Generation

Countering Online Hate Speech: An NLP Perspective

2021-09-07 · Mudit Chaudhary, Chandni Saxena, Helen Meng

Online hate speech has caught everyone's attention from the news related to the COVID-19 pandemic, US elections, and worldwide protests. Online toxicity - an umbrella term for online hateful behavior, manifests itself in…

Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues

2024-01-15 · Sougata Saha, Rohini Srihari

Hateful comments are prevalent on social media platforms. Although tools for automatically detecting, flagging, and blocking such false, offensive, and harmful content online have lately matured, such reactive and brute …

BlockingResponse Generation

Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering

2024-10-04 · Helena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio 외

The potential effectiveness of counterspeech as a hate speech mitigation strategy is attracting increasing interest in the NLG research community, particularly towards the task of automatically producing it. However, aut…