paper-with-me

홈 › Papers

Towards Procedural Fairness: Uncovering Biases in How a Toxic Language Classifier Uses Sentiment Information

2022-10-19 · Isar Nejadgholi, Esma Balkir, Kathleen C. Fraser, Svetlana Kiritchenko

Previous works on the fairness of toxic language classifiers compare the output of models with different identity terms as input features but do not consider the impact of other important concepts present in the context. Here, besides identity terms, we take into account high-level latent features learned by the classifier and investigate the interaction between these features and identity terms. For a multi-class toxic language classifier, we leverage a concept-based explanation framework to calculate the sensitivity of the model to the concept of sentiment, which has been used before as a salient feature for toxic language detection. Our results show that although for some classes, the classifier has learned the sentiment information as expected, this information is outweighed by the influence of identity terms as input features. This work is a step towards evaluating procedural fairness, where unfair processes lead to unfair outcomes. The produced knowledge can guide debiasing techniques to ensure that important concepts besides identity terms are well-represented in training datasets.

📄 PDF Abstract BibTeX arXiv:2210.10689

Code (1)

isarnejad/procedural-fairness-sentiment 공식 구현 pytorch

Tasks

Fairness

Similar Papers 제목 키워드 기반

Addressing Biases in the Texts using an End-to-End Pipeline Approach

2023-03-13 · Shaina Raza, Syed Raza Bashir, Sneha, Urooj Qamar

The concept of fairness is gaining popularity in academia and industry. Social media is especially vulnerable to media biases and toxic language and comments. We propose a fair ML pipeline that takes a text as input and …

FairnessWord Embeddings

On Bias and Fairness in NLP: Investigating the Impact of Bias and Debiasing in Language Models on the Fairness of Toxicity Detection

2023-05-22 · Fatma Elsafoury, Stamos Katsigiannis

Language models are the new state-of-the-art natural language processing (NLP) models and they are being increasingly used in many NLP tasks. Even though there is evidence that language models are biased, the impact of t…

ClassificationFairnessSelection biastext-classification+1

Mitigating Racial Biases in Toxic Language Detection with an Equity-Based Ensemble Framework

2021-09-27 · Matan Halevy, Camille Harris, Amy Bruckman, Diyi Yang 외

Recent research has demonstrated how racial biases against users who write African American English exists in popular toxic language datasets. While previous work has focused on a single fairness criteria, we propose to …

DescriptiveFairness

Procedural Fairness and Its Relationship with Distributive Fairness in Machine Learning

2025-01-12 · ZiMing Wang, Changwu Huang, Ke Tang, Xin Yao

Fairness in machine learning (ML) has garnered significant attention in recent years. While existing research has predominantly focused on the distributive fairness of ML models, there has been limited exploration of pro…

Decision MakingFairness

When to Invoke: Refining LLM Fairness with Toxicity Assessment

2026-01-14 · Jing Ren, Bowen Li, Ziqi Xu, Renqiang Luo 외 arxiv

Large Language Models (LLMs) are increasingly used for toxicity assessment in online moderation systems, where fairness across demographic groups is essential for equitable treatment. However, LLMs often produce inconsis…