Beyond Detection: Unveiling Fairness Vulnerabilities in Abusive Language Models
This work investigates the potential of undermining both fairness and detection performance in abusive language detection. In a dynamic and complex digital world, it is crucial to investigate the vulnerabilities of these detection models to adversarial fairness attacks to improve their fairness robustness. We propose a simple yet effective framework FABLE that leverages backdoor attacks as they allow targeted control over the fairness and detection performance. FABLE explores three types of trigger designs (i.e., rare, artificial, and natural triggers) and novel sampling strategies. Specifically, the adversary can inject triggers into samples in the minority group with the favored outcome (i.e., "non-abusive") and flip their labels to the unfavored outcome, i.e., "abusive". Experiments on benchmark datasets demonstrate the effectiveness of FABLE attacking fairness and utility in abusive language detection.
Code (0)
등록된 구현이 없습니다.
Tasks
Abusive LanguageFairnessMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Fine-Grained Fairness Analysis of Abusive Language Detection Systems with CheckList
Current abusive language detection systems have demonstrated unintended bias towards sensitive features such as nationality or gender. This is a crucial issue, which may harm minorities and underrepresented groups if suc…
Abusive LanguageFairnessBeyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
Multimodal Retrieval-Augmented Generation (MRAG) systems enhance LMMs by integrating external multimodal databases, but introduce unexplored privacy vulnerabilities. While text-based RAG privacy risks have been studied, …
Privacy PreservingRAGRetrievalRetrieval-augmented GenerationAbusive and Threatening Language Detection in Urdu using Boosting based and BERT based models: A Comparative Approach
Online hatred is a growing concern on many social media platforms. To address this issue, different social media platforms have introduced moderation policies for such content. They also employ moderators who can check t…
Abusive LanguageAbusive Language Detection and Characterization of Twitter Behavior
In this work, abusive language detection in online content is performed using Bidirectional Recurrent Neural Network (BiRNN) method. Here the main objective is to focus on various forms of abusive behaviors on Twitter an…
Abusive LanguageMUCIC@TamilNLP-ACL2022: Abusive Comment Detection in Tamil Language using 1D Conv-LSTM
Abusive language content such as hate speech, profanity, and cyberbullying etc., which is common in online platforms is creating lot of problems to the users as well as policy makers. Hence, detection of such abusive lan…
Abusive Language