paper-with-me

Papers

Towards Fairness Assessment of Dutch Hate Speech Detection

2025-06-14 · Julie Bauer, Rishabh Kaushal, Thales Bertaglia, Adriana Iamnitchi

Numerous studies have proposed computational methods to detect hate speech online, yet most focus on the English language and emphasize model development. In this study, we evaluate the counterfactual fairness of hate speech detection models in the Dutch language, specifically examining the performance and fairness of transformer-based models. We make the following key contributions. First, we curate a list of Dutch Social Group Terms that reflect social context. Second, we generate counterfactual data for Dutch hate speech using LLMs and established strategies like Manual Group Substitution (MGS) and Sentence Log-Likelihood (SLL). Through qualitative evaluation, we highlight the challenges of generating realistic counterfactuals, particularly with Dutch grammar and contextual coherence. Third, we fine-tune baseline transformer-based models with counterfactual data and evaluate their performance in detecting hate speech. Fourth, we assess the fairness of these models using Counterfactual Token Fairness (CTF) and group fairness metrics, including equality of odds and demographic parity. Our analysis shows that models perform better in terms of hate speech detection, average counterfactual fairness and group fairness. This work addresses a significant gap in the literature on counterfactual fairness for hate speech detection in Dutch and provides practical insights and recommendations for improving both model performance and fairness.

📄 PDF Abstract BibTeX arXiv:2506.12502

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualFairnessHate Speech DetectionSentence

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Improving Hate Speech Type and Target Detection with Hateful Metaphor Features

2021-06-01 · NAACL (NLP4IF) 2021 6 · Jens Lemmens, Ilia Markov, Walter Daelemans

We study the usefulness of hateful metaphorsas features for the identification of the type and target of hate speech in Dutch Facebook comments. For this purpose, all hateful metaphors in the Dutch LiLaH corpus were anno…

Vocal Bursts Type Prediction

The Role of Context in Detecting the Target of Hate Speech

2022-10-01 · TRAC (COLING) 2022 10 · Ilia Markov, Walter Daelemans

Online hate speech detection is an inherently challenging task that has recently received much attention from the natural language processing community. Despite a substantial increase in performance, considerable challen…

Hate Speech DetectionLanguage ModelingLanguage Modelling

Features or Spurious Artifacts? Data-centric Baselines for Fair and Robust Hate Speech Detection

2022-07-01 · NAACL 2022 7 · Alan Ramponi, Sara Tonelli

Avoiding to rely on dataset artifacts to predict hate speech is at the cornerstone of robust and fair hate speech detection. In this paper we critically analyze lexical biases in hate speech detection via a cross-platfor…

FairnessHate Speech Detection

Exploring Stylometric and Emotion-Based Features for Multilingual Cross-Domain Hate Speech Detection

2021-04-01 · EACL (WASSA) 2021 4 · Ilia Markov, Nikola Ljubešić, Darja Fišer, Walter Daelemans

In this paper, we describe experiments designed to evaluate the impact of stylometric and emotion-based features on hate speech detection: the task of classifying textual content into hate or non-hate speech classes. Our…

Hate Speech Detection

Aligning Attention with Human Rationales for Self-Explaining Hate Speech Detection

2025-11-10 · Brage Eilertsen, Røskva Bjørgfinsdóttir, Francielle Vargas, Ali Ramezani-Kebrya arxiv

The opaque nature of deep learning models presents significant challenges for the ethical deployment of hate speech detection systems. To address this limitation, we introduce Supervised Rational Attention (SRA), a frame…

Hate Speech Detection