Using Pre-Trained Language Models for Producing Counter Narratives Against Hate Speech: a Comparative Study
In this work, we present an extensive study on the use of pre-trained language models for the task of automatic Counter Narrative (CN) generation to fight online hate speech in English. We first present a comparative study to determine whether there is a particular Language Model (or class of LMs) and a particular decoding mechanism that are the most appropriate to generate CNs. Findings show that autoregressive models combined with stochastic decodings are the most promising. We then investigate how an LM performs in generating a CN with regard to an unseen target of hate. We find out that a key element for successful `out of target' experiments is not an overall similarity with the training data but the presence of a specific subset of training data, i.e. a target that shares some commonalities with the test target that can be defined a-priori. We finally introduce the idea of a pipeline based on the addition of an automatic post-editing step to refine generated CNs.
Code (1)
Tasks
Automatic Post-EditingLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Parsimonious Argument Annotations for Hate Speech Counter-narratives
We present an enrichment of the Hateval corpus of hate speech tweets (Basile et. al 2019) aimed to facilitate automated counter-narrative generation. Comparably to previous work (Chung et. al. 2019), manually written cou…
Weigh Your Own Words: Improving Hate Speech Counter Narrative Generation via Attention Regularization
Recent computational approaches for combating online hate speech involve the automatic generation of counter narratives by adapting Pretrained Transformer-based Language Models (PLMs) with human-curated data. This proces…
A feast for trolls -- Engagement analysis of counternarratives against online toxicity
This report provides an engagement analysis of counternarratives against online toxicity. Between February 2020 and July 2021, we observed over 15 million toxic messages on social media identified by our fine-grained, mu…
Multilingual Counter Narrative Type Classification
The growing interest in employing counter narratives for hatred intervention brings with it a focus on dataset creation and automation strategies. In this scenario, learning to recognize counter narrative types from natu…
ClassificationVocal Bursts Type PredictionUsing Instruction-Tuned Large Language Models to Identify Indicators of Vulnerability in Police Incident Narratives
Objectives: Compare qualitative coding of instruction tuned large language models (IT-LLMs) against human coders in classifying the presence or absence of vulnerability in routinely collected unstructured text that descr…
counterfactualGeneral ClassificationLLM real-life tasksSpecificity