paper-with-me

Papers

HateCheck: Functional Tests for Hate Speech Detection Models

2020-12-31 · ACL 2021 5 · Paul Röttger, Bertram Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, Janet B. Pierrehumbert

Detecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model performance due to increasingly well-evidenced systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, we introduce HateCheck, a suite of functional tests for hate speech detection models. We specify 29 model functionalities motivated by a review of previous research and a series of interviews with civil society stakeholders. We craft test cases for each functionality and validate their quality through a structured annotation process. To illustrate HateCheck's utility, we test near-state-of-the-art transformer models as well as two popular commercial models, revealing critical model weaknesses.

📄 PDF Abstract BibTeX arXiv:2012.15606

Code (3)

paul-rottger/hate-functional-tests 공식 구현
paul-rottger/hatecheck-data 공식 구현
paul-rottger/hatecheck-experiments 공식 구현

Tasks

DiagnosticHate Speech Detection

Similar Papers 제목 키워드 기반

SEAHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Southeast Asia

2026-03-17 · Ri Chi Ng, Aditi Kumaresan, Yujia Hu, Roy Ka-Wei Lee arxiv

Hate speech detection relies heavily on linguistic resources, which are primarily available in high-resource languages such as English and Chinese, creating barriers for researchers and platforms developing tools for low…

Hate Speech Detection

Multilingual HateCheck: Functional Tests for Multilingual Hate Speech Detection Models

2022-06-20 · NAACL (WOAH) 2022 7 · Paul Röttger, Haitham Seelawi, Debora Nozza, Zeerak Talat 외

Hate speech detection models are typically evaluated on held-out test sets. However, this risks painting an incomplete and potentially misleading picture of model performance because of increasingly well-documented syste…

DiagnosticHate Speech Detection

SGHateCheck: Functional Tests for Detecting Hate Speech in Low-Resource Languages of Singapore

2024-05-03 · Ri Chi Ng, Nirmalendu Prakash, Ming Shan Hee, Kenny Tsu Wei Choo 외

To address the limitations of current hate speech detection models, we introduce \textsf{SGHateCheck}, a novel framework designed for the linguistic and cultural context of Singapore and Southeast Asia. It extends the fu…

Hate Speech DetectionTranslation

GPT-HateCheck: Can LLMs Write Better Functional Tests for Hate Speech Detection?

2024-02-23 · Yiping Jin, Leo Wanner, Alexander Shvets

Online hate detection suffers from biases incurred in data sampling, annotation, and model pre-training. Therefore, measuring the averaged performance over all examples in held-out test data is inadequate. Instead, we mu…

DiagnosticHate Speech DetectionNatural Language InferenceSentence

Checking HateCheck: a cross-functional analysis of behaviour-aware learning for hate speech detection

2022-04-08 · nlppower (ACL) 2022 5 · Pedro Henrique Luz de Araujo, Benjamin Roth

Behavioural testing -- verifying system capabilities by validating human-designed input-output pairs -- is an alternative evaluation method of natural language processing systems proposed to address the shortcomings of t…

Hate Speech Detection