paper-with-me

홈 › Papers

Prediction Uncertainty Estimation for Hate Speech Classification

2019-09-16 · Kristian Miok, Dong Nguyen-Doan, Blaž Škrlj, Daniela Zaharie, Marko Robnik-Šikonja

As a result of social network popularity, in recent years, hate speech phenomenon has significantly increased. Due to its harmful effect on minority groups as well as on large communities, there is a pressing need for hate speech detection and filtering. However, automatic approaches shall not jeopardize free speech, so they shall accompany their decisions with explanations and assessment of uncertainty. Thus, there is a need for predictive machine learning models that not only detect hate speech but also help users understand when texts cross the line and become unacceptable. The reliability of predictions is usually not addressed in text classification. We fill this gap by proposing the adaptation of deep neural networks that can efficiently estimate prediction uncertainty. To reliably detect hate speech, we use Monte Carlo dropout regularization, which mimics Bayesian inference within neural networks. We evaluate our approach using different text embedding methods. We visualize the reliability of results with a novel technique that aids in understanding the classification reliability and errors.

📄 PDF Abstract BibTeX arXiv:1909.07158

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian InferenceClassificationGeneral ClassificationHate Speech DetectionPredictiontext-classificationText Classification

Methods 이 논문이 사용한 방법론

Monte Carlo Dropout 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection

2024-02-18 · Min Zhang, Jianfeng He, Taoran Ji, Chang-Tien Lu

The fairness and trustworthiness of Large Language Models (LLMs) are receiving increasing attention. Implicit hate speech, which employs indirect language to convey hateful intentions, occupies a significant portion of p…

FairnessHate Speech DetectionSensitivity

To BAN or not to BAN: Bayesian Attention Networks for Reliable Hate Speech Detection

2020-07-10 · Kristian Miok, Blaz Skrlj, Daniela Zaharie, Marko Robnik-Sikonja

Hate speech is an important problem in the management of user-generated content. To remove offensive content or ban misbehaving users, content moderators need reliable hate speech detectors. Recently, deep neural network…

ClassificationGeneral ClassificationHate Speech DetectionManagement

Hierarchical CVAE for Fine-Grained Hate Speech Classification

2018-08-31 · EMNLP 2018 10 · Jing Qian, Mai ElSherief, Elizabeth Belding, William Yang Wang

Existing work on automated hate speech detection typically focuses on binary classification or on differentiating among a small set of categories. In this paper, we propose a novel method on a fine-grained hate speech cl…

Binary ClassificationClassificationGeneral ClassificationHate Speech Detection

An Effective, Robust and Fairness-aware Hate Speech Detection Framework

2024-09-25 · Guanyi Mou, Kyumin Lee

With the widespread online social networks, hate speeches are spreading faster and causing more damage than ever before. Existing hate speech detection methods have limitations in several aspects, such as handling data i…

FairnessHate Speech Detection

Hypothesis Engineering for Zero-Shot Hate Speech Detection

2022-10-03 · TRAC (COLING) 2022 10 · Janis Goldzycher, Gerold Schneider

Standard approaches to hate speech detection rely on sufficient available hate speech annotations. Extending previous work that repurposes natural language inference (NLI) models for zero-shot text classification, we pro…

Hate Speech DetectionNatural Language InferenceText ClassificationZero-Shot Text Classification