paper-with-me

홈 › Papers

Evaluating Large Language Models for Detecting Antisemitism

2025-09-22 · Jay Patel, Hrudayangam Mehta, Jeremy Blackburn arxiv

Detecting hateful content is a challenging and important problem. Automated tools, like machine-learning models, can help, but they require continuous training to adapt to the ever-changing landscape of social media. In this work, we evaluate eight open-source LLMs' capability to detect antisemitic content, specifically leveraging in-context definition. We also study how LLMs understand and explain their decisions given a moderation policy as a guideline. First, we explore various prompting techniques and design a new CoT-like prompt, Guided-CoT, and find that injecting domain-specific thoughts increases performance and utility. Guided-CoT handles the in-context policy well, improving performance and utility by reducing refusals across all evaluated models, regardless of decoding configuration, model size, or reasoning capability. Notably, Llama 3.1 70B outperforms fine-tuned GPT-3.5. Additionally, we examine LLM errors and introduce metrics to quantify semantic divergence in model-generated rationales, revealing notable differences and paradoxical behaviors among LLMs. Our experiments highlight the differences observed across LLMs' utility, explainability, and reliability. Code and resources available at: https://github.com/idramalab/quantify-llm-explanations

📄 PDF Abstract BibTeX arXiv:2509.18293

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

You Frame It: How Conceptual Representations Shape LLM Detection and Reasoning about Antisemitism

2026-07-06 · Katharina Soemer, Helena Mihaljević arxiv

LLMs enable the integration of external conceptual resources at inference time, creating new opportunities for detecting ideologically and historically complex phenomena such as antisemitism. We investigate how different…

Tracing Antisemitic Language Through Diachronic Embedding Projections: France 1789-1914

2019-06-04 · WS 2019 8 · Rocco Tripodi, Massimo Warglien, Simon Levis Sullam, Deborah Paci

We investigate some aspects of the history of antisemitism in France, one of the cradles of modern antisemitism, using diachronic word embeddings. We constructed a large corpus of French books and periodicals issues that…

Diachronic Word EmbeddingsWord Embeddings

"Subverting the Jewtocracy": Online Antisemitism Detection Using Multimodal Deep Learning

2021-04-13 · Mohit Chandra, Dheeraj Pailla, Himanshu Bhatia, AadilMehdi Sanchawala 외

The exponential rise of online social media has enabled the creation, distribution, and consumption of information at an unprecedented rate. However, it has also led to the burgeoning of various forms of online abuse. In…

Deep LearningMultimodal Deep Learning

Codes, Patterns and Shapes of Contemporary Online Antisemitism and Conspiracy Narratives -- an Annotation Guide and Labeled German-Language Dataset in the Context of COVID-19

2022-10-13 · Elisabeth Steffen, Helena Mihaljević, Milena Pustet, Nyco Bischoff 외

Over the course of the COVID-19 pandemic, existing conspiracy theories were refreshed and new ones were created, often interwoven with antisemitic narratives, stereotypes and codes. The sheer volume of antisemitic and co…

Antisemitic Messages? A Guide to High-Quality Annotation and a Labeled Dataset of Tweets

2023-04-28 · Gunther Jikeli, Sameer Karali, Daniel Miehling, Katharina Soemer

One of the major challenges in automatic hate speech detection is the lack of datasets that cover a wide range of biased and unbiased messages and that are consistently labeled. We propose a labeling procedure that addre…

Hate Speech Detection