paper-with-me

홈 › Papers

Sociocultural knowledge is needed for selection of shots in hate speech detection tasks

2023-04-04 · Antonis Maronikolakis, Abdullatif Köksal, Hinrich Schütze

We introduce HATELEXICON, a lexicon of slurs and targets of hate speech for the countries of Brazil, Germany, India and Kenya, to aid training and interpretability of models. We demonstrate how our lexicon can be used to interpret model predictions, showing that models developed to classify extreme speech rely heavily on target words when making predictions. Further, we propose a method to aid shot selection for training in low-resource settings via HATELEXICON. In few-shot learning, the selection of shots is of paramount importance to model performance. In our work, we simulate a few-shot setting for German and Hindi, using HASOC data for training and the Multilingual HateCheck (MHC) as a benchmark. We show that selecting shots based on our lexicon leads to models performing better on MHC than models trained on shots sampled randomly. Thus, when given only a few training examples, using our lexicon to select shots containing more sociocultural information leads to better few-shot performance.

📄 PDF Abstract BibTeX arXiv:2304.01890

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot LearningHate Speech Detection

Similar Papers 제목 키워드 기반

Sociocultural Considerations in Monitoring Anti-LGBTQ+ Content on Social Media

2024-07-01 · Sidney G. -J. Wong

The purpose of this paper is to ascertain the influence of sociocultural factors (i.e., social, cultural, and political) in the development of hate speech detection systems. We set out to investigate the suitability of u…

Hate Speech Detection

SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant Hostility

2026-01-28 · Xuanyu Su, Diana Inkpen, Nathalie Japkowicz arxiv

Online hate on social media ranges from overt slurs and threats (\emph{hard hate speech}) to \emph{soft hate speech}: discourse that appears reasonable on the surface but uses framing and value-based arguments to steer a…

Selecting and combining complementary feature representations and classifiers for hate speech detection

2022-01-18 · Rafael M. O. Cruz, Woshington V. de Sousa, George D. C. Cavalcanti

Hate speech is a major issue in social networks due to the high volume of data generated daily. Recent works demonstrate the usefulness of machine learning (ML) in dealing with the nuances required to distinguish between…

ClassificationHate Speech Detection

Memes in the Wild: Assessing the Generalizability of the Hateful Memes Challenge Dataset

2021-07-09 · ACL (WOAH) 2021 8 · Hannah Rose Kirk, Yennie Jun, Paulius Rauba, Gal Wachtel 외

Hateful memes pose a unique challenge for current machine learning systems because their message is derived from both text- and visual-modalities. To this effect, Facebook released the Hateful Memes Challenge, a dataset …

Optical Character Recognition (OCR)

OntoSOC: Sociocultural Knowledge Ontology

2015-05-15 · Guidedi Kaladzavi, Papa Fary Diallo, Kolyang, Moussa Lo

This paper presents a sociocultural knowledge ontology (OntoSOC) modeling approach. OntoSOC modeling approach is based on Engestrom Human Activity Theory (HAT). That Theory allowed us to identify fundamental concepts and…

Information RetrievalRetrieval