paper-with-me

Papers

Counterfactual Multi-Token Fairness in Text Classification

2022-02-08 · Pranay Lohia

The counterfactual token generation has been limited to perturbing only a single token in texts that are generally short and single sentences. These tokens are often associated with one of many sensitive attributes. With limited counterfactuals generated, the goal to achieve invariant nature for machine learning classification models towards any sensitive attribute gets bounded, and the formulation of Counterfactual Fairness gets narrowed. In this paper, we overcome these limitations by solving root problems and opening bigger domains for understanding. We have curated a resource of sensitive tokens and their corresponding perturbation tokens, even extending the support beyond traditionally used sensitive attributes like Age, Gender, Race to Nationality, Disability, and Religion. The concept of Counterfactual Generation has been extended to multi-token support valid over all forms of texts and documents. We define the method of generating counterfactuals by perturbing multiple sensitive tokens as Counterfactual Multi-token Generation. The method has been conceptualized to showcase significant performance improvement over single-token methods and validated over multiple benchmark datasets. The emendation in counterfactual generation propagates in achieving improved Counterfactual Multi-token Fairness.

📄 PDF Abstract BibTeX arXiv:2202.03792

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeClassificationcounterfactualFairnesstext-classificationText Classificationvalid

Methods 이 논문이 사용한 방법론

Counterfactuals 설명 없음

Similar Papers 제목 키워드 기반

Counterfactual Fairness in Text Classification through Robustness

2018-09-27 · Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly 외

In this paper, we study counterfactual fairness in text classification, which asks the question: How would the prediction change if the sensitive attribute referenced in the example were different? Toxicity classifiers d…

AttributeClassificationcounterfactualFairness+3

Fairness for Text Classification Tasks with Identity Information Data Augmentation Methods

2022-02-04 · Mohit Wadhwa, Mohan Bhambhani, Ashvini Jindal, Uma Sawant 외

Counterfactual fairness methods address the question: How would the prediction change if the sensitive identity attributes referenced in the text instance were different? These methods are entirely based on generating co…

counterfactualData AugmentationFairnesstext-classification+2

Towards Fairness Assessment of Dutch Hate Speech Detection

2025-06-14 · Julie Bauer, Rishabh Kaushal, Thales Bertaglia, Adriana Iamnitchi

Numerous studies have proposed computational methods to detect hate speech online, yet most focus on the English language and emphasize model development. In this study, we evaluate the counterfactual fairness of hate sp…

counterfactualFairnessHate Speech DetectionSentence

Fair Hate Speech Detection through Evaluation of Social Group Counterfactuals

2020-10-24 · Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy, Mohammad Atari 외

Approaches for mitigating bias in supervised models are designed to reduce models' dependence on specific sensitive features of the input data, e.g., mentioned social groups. However, in the case of hate speech detection…

counterfactualFairnessHate Speech DetectionSentence

Improving Counterfactual Generation for Fair Hate Speech Detection

2021-08-03 · ACL (WOAH) 2021 8 · Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy, Mohammad Atari 외

Bias mitigation approaches reduce models' dependence on sensitive features of data, such as social group tokens (SGTs), resulting in equal predictions across the sensitive features. In hate speech detection, however, equ…

counterfactualFairnessHate Speech DetectionSentence