paper-with-me

홈 › Papers

NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps

2024-04-02 · Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Durmus, Dan Jurafsky

The use of words to convey speaker's intent is traditionally distinguished from the `mention' of words for quoting what someone said, or pointing out properties of a word. Here we show that computationally modeling this use-mention distinction is crucial for dealing with counterspeech online. Counterspeech that refutes problematic content often mentions harmful language but is not harmful itself (e.g., calling a vaccine dangerous is not the same as expressing disapproval of someone for calling vaccines dangerous). We show that even recent language models fail at distinguishing use from mention, and that this failure propagates to two key downstream tasks: misinformation and hate speech detection, resulting in censorship of counterspeech. We introduce prompting mitigations that teach the use-mention distinction, and show they reduce these errors. Our work highlights the importance of the use-mention distinction for NLP and CSS and offers ways to address it.

📄 PDF Abstract BibTeX arXiv:2404.01651

Code (1)

kristinagligoric/use-mention 공식 구현

Tasks

Hate Speech DetectionMisinformation

Similar Papers 제목 키워드 기반

Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF

2024-03-15 · Amey Hengle, Aswini Kumar, Sahajpreet Singh, Anil Bandhakavi 외

Counterspeech, defined as a response to mitigate online hate speech, is increasingly used as a non-censorial solution. Addressing hate speech effectively involves dispelling the stereotypes, prejudices, and biases often …

Sentence

Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language

2023-10-31 · Jimin Mun, Emily Allaway, Akhila Yerukola, Laura Vianna 외

Counterspeech, i.e., responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship. However, properly countering hateful language …

Philosophy

Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech Generation

2023-05-23 · Rishabh Gupta, Shaily Desai, Manvi Goel, Anil Bandhakavi 외

Counterspeech has been demonstrated to be an efficacious approach for combating hate speech. While various conventional and controlled approaches have been studied in recent years to generate counterspeech, a counterspee…

Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning

2025-05-17 · Aswini Kumar Padhi, Anil Bandhakavi, Tanmoy Chakraborty

Counterspeech has proven to be a powerful tool to combat hate speech online. Previous studies have focused on generating counterspeech conditioned only on specific intents (single attributed). However, a holistic approac…

Attribute

CounterQuill: Investigating the Potential of Human-AI Collaboration in Online Counterspeech Writing

2024-10-03 · Xiaohan Ding, Kaike Ping, Uma Sushmitha Gunturi, Buse Carik 외

Online hate speech has become increasingly prevalent on social media, causing harm to individuals and society. While automated content moderation has received considerable attention, user-driven counterspeech remains a l…