paper-with-me

Papers

Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language

2023-10-31 · Jimin Mun, Emily Allaway, Akhila Yerukola, Laura Vianna, Sarah-Jane Leslie, Maarten Sap

Counterspeech, i.e., responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship. However, properly countering hateful language requires countering and dispelling the underlying inaccurate stereotypes implied by such language. In this work, we draw from psychology and philosophy literature to craft six psychologically inspired strategies to challenge the underlying stereotypical implications of hateful language. We first examine the convincingness of each of these strategies through a user study, and then compare their usages in both human- and machine-generated counterspeech datasets. Our results show that human-written counterspeech uses countering strategies that are more specific to the implied stereotype (e.g., counter examples to the stereotype, external factors about the stereotype's origins), whereas machine-generated counterspeech uses less specific strategies (e.g., generally denouncing the hatefulness of speech). Furthermore, machine-generated counterspeech often employs strategies that humans deem less convincing compared to human-produced counterspeech. Our findings point to the importance of accounting for the underlying stereotypical implications of speech when generating counterspeech and for better machine reasoning about anti-stereotypical examples.

📄 PDF Abstract BibTeX arXiv:2311.00161

Code (0)

등록된 구현이 없습니다.

Tasks

Philosophy

Similar Papers 제목 키워드 기반

Countering Online Hate Speech: An NLP Perspective

2021-09-07 · Mudit Chaudhary, Chandni Saxena, Helen Meng

Online hate speech has caught everyone's attention from the news related to the COVID-19 pandemic, US elections, and worldwide protests. Online toxicity - an umbrella term for online hateful behavior, manifests itself in…

Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues

2024-01-15 · Sougata Saha, Rohini Srihari

Hateful comments are prevalent on social media platforms. Although tools for automatically detecting, flagging, and blocking such false, offensive, and harmful content online have lately matured, such reactive and brute …

BlockingResponse Generation

A Benchmark Dataset for Learning to Intervene in Online Hate Speech

2019-09-10 · IJCNLP 2019 11 · Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding 외

Countering online hate speech is a critical yet challenging task, but one which can be aided by the use of Natural Language Processing (NLP) techniques. Previous research has primarily focused on the development of NLP m…

Response Generation

Empowering NGOs in Countering Online Hate Messages

2021-07-06 · Yi-Ling Chung, Serra Sinem Tekiroglu, Sara Tonelli, Marco Guerini

Studies on online hate speech have mostly focused on the automated detection of harmful messages. Little attention has been devoted so far to the development of effective strategies to fight hate speech, in particular th…

Management

Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate

2025-06-04 · Mikel K. Ngueajio, Flor Miriam Plaza-del-Arco, Yi-Ling Chung, Danda B. Rawat 외

Automated counter-narratives (CN) offer a promising strategy for mitigating online hate speech, yet concerns about their affective tone, accessibility, and ethical risks remain. We propose a framework for evaluating Larg…

Language ModelingLanguage ModellingLarge Language Model