paper-with-me

Papers

Causality Guided Representation Learning for Cross-Style Hate Speech Detection

2025-10-09 · Chengshuai Zhao, Shu Wan, Paras Sheth, Karan Patwa, K. Selçuk Candan, Huan Liu arxiv

The proliferation of online hate speech poses a significant threat to the harmony of the web. While explicit hate is easily recognized through overt slurs, implicit hate speech is often conveyed through sarcasm, irony, stereotypes, or coded language -- making it harder to detect. Existing hate speech detection models, which predominantly rely on surface-level linguistic cues, fail to generalize effectively across diverse stylistic variations. Moreover, hate speech spread on different platforms often targets distinct groups and adopts unique styles, potentially inducing spurious correlations between them and labels, further challenging current detection approaches. Motivated by these observations, we hypothesize that the generation of hate speech can be modeled as a causal graph involving key factors: contextual environment, creator motivation, target, and style. Guided by this graph, we propose CADET, a causal representation learning framework that disentangles hate speech into interpretable latent factors and then controls confounders, thereby isolating genuine hate intent from superficial linguistic cues. Furthermore, CADET allows counterfactual reasoning by intervening on style within the latent space, naturally guiding the model to robustly identify hate speech in varying forms. CADET demonstrates superior performance in comprehensive experiments, highlighting the potential of causal priors in advancing generalizable hate speech detection.

📄 PDF Abstract BibTeX arXiv:2510.07707

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningHate Speech Detection

Similar Papers 제목 키워드 기반

PEACE: Cross-Platform Hate Speech Detection- A Causality-guided Framework

2023-06-15 · Paras Sheth, Tharindu Kumarage, Raha Moraffah, Aman Chadha 외

Hate speech detection refers to the task of detecting hateful content that aims at denigrating an individual or a group based on their religion, gender, sexual orientation, or other characteristics. Due to the different …

Hate Speech Detection

Causality Guided Disentanglement for Cross-Platform Hate Speech Detection

2023-08-03 · Paras Sheth, Tharindu Kumarage, Raha Moraffah, Aman Chadha 외

Social media platforms, despite their value in promoting open discourse, are often exploited to spread harmful content. Current deep learning and natural language processing models used for detecting this harmful content…

DisentanglementHate Speech Detection

Cross-Platform Hate Speech Detection with Weakly Supervised Causal Disentanglement

2024-04-17 · Paras Sheth, Tharindu Kumarage, Raha Moraffah, Aman Chadha 외

Content moderation faces a challenging task as social media's ability to spread hate speech contrasts with its role in promoting global connectivity. With rapidly evolving slang and hate speech, the adaptability of conve…

DisentanglementHate Speech Detection

CausalDisenSeg: A Causality-Guided Disentanglement Framework with Counterfactual Reasoning for Robust Brain Tumor Segmentation Under Missing Modalities

2026-04-15 · Bo Liu, Yulong Zou, Jin Hong arxiv

In clinical practice, the robustness of deep learning models for multimodal brain tumor segmentation is severely compromised by incomplete MRI data. This vulnerability stems primarily from modality bias, where models exp…

Brain Tumor Segmentation

Reducing Target Group Bias in Hate Speech Detectors

2021-12-07 · Darsh J Shah, Sinong Wang, Han Fang, Hao Ma 외

The ubiquity of offensive and hateful content on online fora necessitates the need for automatic solutions that detect such content competently across target groups. In this paper we show that text classification models …

text-classificationText Classification