Thesis Distillation: Investigating The Impact of Bias in NLP Models on Hate Speech Detection
This paper is a summary of the work done in my PhD thesis. Where I investigate the impact of bias in NLP models on the task of hate speech detection from three perspectives: explainability, offensive stereotyping bias, and fairness. Then, I discuss the main takeaways from my thesis and how they can benefit the broader NLP community. Finally, I discuss important future research directions. The findings of my thesis suggest that the bias in NLP models impacts the task of hate speech detection from all three perspectives. And that unless we start incorporating social sciences in studying bias in NLP models, we will not effectively overcome the current limitations of measuring and mitigating bias in NLP models.
Code (0)
등록된 구현이 없습니다.
Tasks
FairnessHate Speech DetectionSimilar Papers 제목 키워드 기반
Darkness can not drive out darkness: Investigating Bias in Hate SpeechDetection Models
It has become crucial to develop tools for automated hate speech and abuse detection. These tools would help to stop the bullies and the haters and provide a safer environment for individuals especially from marginalized…
Abuse DetectionHate Speech DetectionHateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
Hate speech detection is a socially sensitive and inherently subjective task, with judgments often varying based on personal traits. While prior work has examined how socio-demographic factors influence annotation, the i…
FairnessHate Speech DetectionImpact of Politically Biased Data on Hate Speech Classification
One challenge that social media platforms are facing nowadays is hate speech. Hence, automatic hate speech detection has been increasingly researched in recent years - in particular with the rise of deep learning. A prob…
ClassificationHate Speech DetectionSystematic Offensive Stereotyping (SOS) Bias in Language Models
In this paper, we propose a new metric to measure the SOS bias in language models (LMs). Then, we validate the SOS bias and investigate the effectiveness of removing it. Finally, we investigate the impact of the SOS bias…
FairnessHate Speech DetectionInvestigating Annotator Bias in Abusive Language Datasets
Nowadays, social media platforms use classification models to cope with hate speech and abusive language. The problem of these models is their vulnerability to bias. A prevalent form of bias in hate speech and abusive la…
Abusive Language