Rethinking Offensive Text Detection as a Multi-Hop Reasoning Problem
We introduce the task of implicit offensive text detection in dialogues, where a statement may have either an offensive or non-offensive interpretation, depending on the listener and context. We argue that reasoning is crucial for understanding this broader class of offensive utterances and release SLIGHT, a dataset to support research on this task. Experiments using the data show that state-of-the-art methods of offense detection perform poorly when asked to detect implicitly offensive statements, achieving only ${\sim} 11\%$ accuracy. In contrast to existing offensive text detection datasets, SLIGHT features human-annotated chains of reasoning which describe the mental process by which an offensive interpretation can be reached from each ambiguous statement. We explore the potential for a multi-hop reasoning approach by utilizing existing entailment models to score the probability of these chains and show that even naive reasoning models can yield improved performance in most situations. Furthermore, analysis of the chains provides insight into the human interpretation process and emphasizes the importance of incorporating additional commonsense knowledge.
Code (1)
Tasks
Text DetectionSimilar Papers 제목 키워드 기반
Rethinking Offensive Text Detection as a Multi-Hop Reasoning Problem
We introduce the task of implicit offensive language detection in dialogues, where a statement may have either an offensive or unoffensive interpretation, depending on the listener and context. We argue that inference is…
Text DetectionLanguage, Culture, and Ideology: Personalizing Offensiveness Detection in Political Tweets with Reasoning LLMs
We explore how large language models (LLMs) assess offensiveness in political discourse when prompted to adopt specific political and cultural perspectives. Using a multilingual subset of the MD-Agreement dataset centere…
Text ClassificationCOBRA Frames: Contextual Reasoning about Effects and Harms of Offensive Statements
Warning: This paper contains content that may be offensive or upsetting. Understanding the harms and offensiveness of statements requires reasoning about the social and situational context in which statements are made. F…
Multimodal Meme Dataset (MultiOFF) for Identifying Offensive Content in Image and Text
A meme is a form of media that spreads an idea or emotion across the internet. As posting meme has become a new form of communication of the web, due to the multimodal nature of memes, postings of hateful memes or relate…
Abuse DetectionMeme ClassificationA multilingual dataset for offensive language and hate speech detection for hausa, yoruba and igbo languages
The proliferation of online offensive language necessitates the development of effective detection mechanisms, especially in multilingual contexts. This study addresses the challenge by developing and introducing novel d…
Hate Speech Detection