paper-with-me

홈 › Papers

Antisemitic Messages? A Guide to High-Quality Annotation and a Labeled Dataset of Tweets

2023-04-28 · Gunther Jikeli, Sameer Karali, Daniel Miehling, Katharina Soemer

One of the major challenges in automatic hate speech detection is the lack of datasets that cover a wide range of biased and unbiased messages and that are consistently labeled. We propose a labeling procedure that addresses some of the common weaknesses of labeled datasets. We focus on antisemitic speech on Twitter and create a labeled dataset of 6,941 tweets that cover a wide range of topics common in conversations about Jews, Israel, and antisemitism between January 2019 and December 2021 by drawing from representative samples with relevant keywords. Our annotation process aims to strictly apply a commonly used definition of antisemitism by forcing annotators to specify which part of the definition applies, and by giving them the option to personally disagree with the definition on a case-by-case basis. Labeling tweets that call out antisemitism, report antisemitism, or are otherwise related to antisemitism (such as the Holocaust) but are not actually antisemitic can help reduce false positives in automated detection. The dataset includes 1,250 tweets (18%) that are antisemitic according to the International Holocaust Remembrance Alliance (IHRA) definition of antisemitism. It is important to note, however, that the dataset is not comprehensive. Many topics are still not covered, and it only includes tweets collected from Twitter between January 2019 and December 2021. Additionally, the dataset only includes tweets that were written in English. Despite these limitations, we hope that this is a meaningful contribution to improving the automated detection of antisemitic speech.

📄 PDF Abstract BibTeX arXiv:2304.14599

Code (0)

등록된 구현이 없습니다.

Tasks

Hate Speech Detection

Similar Papers 제목 키워드 기반

Codes, Patterns and Shapes of Contemporary Online Antisemitism and Conspiracy Narratives -- an Annotation Guide and Labeled German-Language Dataset in the Context of COVID-19

2022-10-13 · Elisabeth Steffen, Helena Mihaljević, Milena Pustet, Nyco Bischoff 외

Over the course of the COVID-19 pandemic, existing conspiracy theories were refreshed and new ones were created, often interwoven with antisemitic narratives, stereotypes and codes. The sheer volume of antisemitic and co…

How toxic is antisemitism? Potentials and limitations of automated toxicity scoring for antisemitic online content

2023-10-05 · Helena Mihaljević, Elisabeth Steffen

The Perspective API, a popular text toxicity assessment service by Google and Jigsaw, has found wide adoption in several application areas, notably content moderation, monitoring, and social media research. We examine it…

Tracing Antisemitic Language Through Diachronic Embedding Projections: France 1789-1914

2019-06-04 · WS 2019 8 · Rocco Tripodi, Massimo Warglien, Simon Levis Sullam, Deborah Paci

We investigate some aspects of the history of antisemitism in France, one of the cradles of modern antisemitism, using diachronic word embeddings. We constructed a large corpus of French books and periodicals issues that…

Diachronic Word EmbeddingsWord Embeddings

Evaluating Large Language Models for Antisemitic Incident Classification

2026-07-06 · Karina Halevy, Julia Mendelsohn, Chan Young Park, Yulia Tsvetkov 외 arxiv

Addressing hate and violence in society requires timely detection of hateful events from public reporting, but automated identification of hateful events remains underexplored. We introduce the task of hateful event dete…

Using LLMs to discover emerging coded antisemitic hate-speech in extremist social media

2024-01-19 · Dhanush Kikkisetti, Raza Ul Mustafa, Wendy Melillo, Roberto Corizzo 외

Online hate speech proliferation has created a difficult problem for social media platforms. A particular challenge relates to the use of coded language by groups interested in both creating a sense of belonging for its …

Language ModelingLanguage ModellingLarge Language ModelSemantic Similarity+1