Prediction Uncertainty Estimation for Hate Speech Classification
As a result of social network popularity, in recent years, hate speech phenomenon has significantly increased. Due to its harmful effect on minority groups as well as on large communities, there is a pressing need for hate speech detection and filtering. However, automatic approaches shall not jeopardize free speech, so they shall accompany their decisions with explanations and assessment of uncertainty. Thus, there is a need for predictive machine learning models that not only detect hate speech but also help users understand when texts cross the line and become unacceptable. The reliability of predictions is usually not addressed in text classification. We fill this gap by proposing the adaptation of deep neural networks that can efficiently estimate prediction uncertainty. To reliably detect hate speech, we use Monte Carlo dropout regularization, which mimics Bayesian inference within neural networks. We evaluate our approach using different text embedding methods. We visualize the reliability of results with a novel technique that aids in understanding the classification reliability and errors.
Code (0)
등록된 구현이 없습니다.
Tasks
Bayesian InferenceClassificationGeneral ClassificationHate Speech DetectionPredictiontext-classificationText ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection
The fairness and trustworthiness of Large Language Models (LLMs) are receiving increasing attention. Implicit hate speech, which employs indirect language to convey hateful intentions, occupies a significant portion of p…
FairnessHate Speech DetectionSensitivityTo BAN or not to BAN: Bayesian Attention Networks for Reliable Hate Speech Detection
Hate speech is an important problem in the management of user-generated content. To remove offensive content or ban misbehaving users, content moderators need reliable hate speech detectors. Recently, deep neural network…
ClassificationGeneral ClassificationHate Speech DetectionManagementHierarchical CVAE for Fine-Grained Hate Speech Classification
Existing work on automated hate speech detection typically focuses on binary classification or on differentiating among a small set of categories. In this paper, we propose a novel method on a fine-grained hate speech cl…
Binary ClassificationClassificationGeneral ClassificationHate Speech DetectionAn Effective, Robust and Fairness-aware Hate Speech Detection Framework
With the widespread online social networks, hate speeches are spreading faster and causing more damage than ever before. Existing hate speech detection methods have limitations in several aspects, such as handling data i…
FairnessHate Speech DetectionHypothesis Engineering for Zero-Shot Hate Speech Detection
Standard approaches to hate speech detection rely on sufficient available hate speech annotations. Extending previous work that repurposes natural language inference (NLI) models for zero-shot text classification, we pro…
Hate Speech DetectionNatural Language InferenceText ClassificationZero-Shot Text Classification