paper-with-me

홈 › Papers

Evaluating Large Language Models for Antisemitic Incident Classification

2026-07-06 · Karina Halevy, Julia Mendelsohn, Chan Young Park, Yulia Tsvetkov, Maarten Sap arxiv

Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We introduce the task of hateful event detection and investigate the ability of AI systems, specifically large language models (LLMs), to discover and classify reports of antisemitic events with fine-grained labels. We evaluate OpenAI's GPT-4o and Meta's Llama-3.2-3B-Instruct on multiple expert-annotated datasets containing antisemitic event descriptions from news articles, civil society reports, and official records. We show that LLMs, particularly GPT-4o, have potential for this task, but substantial improvement is needed. Providing clear term definitions and in-context examples in prompts can improve performance: definitions are most helpful for rhetoric-oriented events (e.g. classical antisemitic tropes), while examples help label action-oriented events (e.g. physical assault). A case study of college newspapers demonstrates that LLMs can help surface relevant real-world events, supporting early monitoring and intervention. Overall, our findings highlight both opportunities and critical gaps in AI's ability to recognize complex harms and underscore the need for collaborative efforts among AI developers, policymakers, and civil society to design models, implement robust evaluation, and develop policy frameworks for defining and combating hate efficiently and effectively.

📄 PDF Abstract BibTeX arXiv:2607.04890

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Codes, Patterns and Shapes of Contemporary Online Antisemitism and Conspiracy Narratives -- an Annotation Guide and Labeled German-Language Dataset in the Context of COVID-19

2022-10-13 · Elisabeth Steffen, Helena Mihaljević, Milena Pustet, Nyco Bischoff 외

Over the course of the COVID-19 pandemic, existing conspiracy theories were refreshed and new ones were created, often interwoven with antisemitic narratives, stereotypes and codes. The sheer volume of antisemitic and co…

Using LLMs to discover emerging coded antisemitic hate-speech in extremist social media

2024-01-19 · Dhanush Kikkisetti, Raza Ul Mustafa, Wendy Melillo, Roberto Corizzo 외

Online hate speech proliferation has created a difficult problem for social media platforms. A particular challenge relates to the use of coded language by groups interested in both creating a sense of belonging for its …

Language ModelingLanguage ModellingLarge Language ModelSemantic Similarity+1

How toxic is antisemitism? Potentials and limitations of automated toxicity scoring for antisemitic online content

2023-10-05 · Helena Mihaljević, Elisabeth Steffen

The Perspective API, a popular text toxicity assessment service by Google and Jigsaw, has found wide adoption in several application areas, notably content moderation, monitoring, and social media research. We examine it…

Tracing Antisemitic Language Through Diachronic Embedding Projections: France 1789-1914

2019-06-04 · WS 2019 8 · Rocco Tripodi, Massimo Warglien, Simon Levis Sullam, Deborah Paci

We investigate some aspects of the history of antisemitism in France, one of the cradles of modern antisemitism, using diachronic word embeddings. We constructed a large corpus of French books and periodicals issues that…

Diachronic Word EmbeddingsWord Embeddings

A Group-Specific Approach to NLP for Hate Speech Detection

2023-04-21 · Karina Halevy

Automatic hate speech detection is an important yet complex task, requiring knowledge of common sense, stereotypes of protected groups, and histories of discrimination, each of which may constantly evolve. In this paper,…

Common Sense ReasoningEthicsHate Speech Detection