paper-with-me

홈 › Papers

Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models

2025-06-10 · Shuzhou Yuan, Ercong Nie, Mario Tawfelis, Helmut Schmid, Hinrich Schütze, Michael Färber

Hate speech detection is a socially sensitive and inherently subjective task, with judgments often varying based on personal traits. While prior work has examined how socio-demographic factors influence annotation, the impact of personality traits on Large Language Models (LLMs) remains largely unexplored. In this paper, we present the first comprehensive study on the role of persona prompts in hate speech classification, focusing on MBTI-based traits. A human annotation survey confirms that MBTI dimensions significantly affect labeling behavior. Extending this to LLMs, we prompt four open-source models with MBTI personas and evaluate their outputs across three hate speech datasets. Our analysis uncovers substantial persona-driven variation, including inconsistencies with ground truth, inter-persona disagreement, and logit-level biases. These findings highlight the need to carefully define persona prompts in LLM-based annotation workflows, with implications for fairness and alignment with human values.

📄 PDF Abstract BibTeX arXiv:2506.08593

Code (0)

등록된 구현이 없습니다.

Tasks

FairnessHate Speech Detection

Similar Papers 제목 키워드 기반

On the Evolution of (Hateful) Memes by Means of Multimodal Contrastive Learning

2022-12-13 · Yiting Qu, Xinlei He, Shannon Pierson, Michael Backes 외

The dissemination of hateful memes online has adverse effects on social media platforms and the real world. Detecting hateful memes is challenging, one of the reasons being the evolutionary nature of memes; new hateful m…

Contrastive Learning

A Web of Hate: Tackling Hateful Speech in Online Social Spaces

2017-09-28 · Haji Mohammad Saleem, Kelly P Dillon, Susan Benesch, Derek Ruths

Online social platforms are beset with hateful speech - content that expresses hatred for a person or group of people. Such content can frighten, intimidate, or silence platform users, and some of it can inspire other us…

Deradicalizing YouTube: Characterization, Detection, and Personalization of Religiously Intolerant Arabic Videos

2022-06-30 · Nuha Albadi, Maram Kurdi, Shivakant Mishra

Growing evidence suggests that YouTube's recommendation algorithm plays a role in online radicalization via surfacing extreme content. Radical Islamist groups, in particular, have been profiting from the global appeal of…

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions

2025-07-30 · Yiting Qu, Ziqing Yang, Yihan Ma, Michael Backes 외 arxiv

Recent advances in text-to-image diffusion models have enabled the creation of a new form of digital art: optical illusions--visual tricks that create different perceptions of reality. However, adversaries may misuse suc…

Improving Hateful Meme Detection through Retrieval-Guided Contrastive Learning

2023-11-14 · Jingbiao Mei, Jinghong Chen, Weizhe Lin, Bill Byrne 외

Hateful memes have emerged as a significant concern on the Internet. Detecting hateful memes requires the system to jointly understand the visual and textual modalities. Our investigation reveals that the embedding space…

Contrastive LearningHateful Meme ClassificationMeme ClassificationRetrieval