paper-with-me

홈 › Papers

Perturbation Sensitivity Analysis to Detect Unintended Model Biases

2019-10-09 · IJCNLP 2019 11 · Vinodkumar Prabhakaran, Ben Hutchinson, Margaret Mitchell

Data-driven statistical Natural Language Processing (NLP) techniques leverage large amounts of language data to build models that can understand language. However, most language data reflect the public discourse at the time the data was produced, and hence NLP models are susceptible to learning incidental associations around named referents at a particular point in time, in addition to general linguistic meaning. An NLP system designed to model notions such as sentiment and toxicity should ideally produce scores that are independent of the identity of such entities mentioned in text and their social associations. For example, in a general purpose sentiment analysis system, a phrase such as I hate Katy Perry should be interpreted as having the same sentiment as I hate Taylor Swift. Based on this idea, we propose a generic evaluation framework, Perturbation Sensitivity Analysis, which detects unintended model biases related to named entities, and requires no new annotations or corpora. We demonstrate the utility of this analysis by employing it on two different NLP models --- a sentiment model and a toxicity model --- applied on online comments in English language from four different genres.

📄 PDF Abstract BibTeX arXiv:1910.04210

Code (0)

등록된 구현이 없습니다.

Tasks

modelSensitivitySentiment Analysis

Similar Papers 제목 키워드 기반

Image Counterfactual Sensitivity Analysis for Detecting Unintended Bias

2019-06-14 · Emily Denton, Ben Hutchinson, Margaret Mitchell, Timnit Gebru 외

Facial analysis models are increasingly used in applications that have serious impacts on people's lives, ranging from authentication to surveillance tracking. It is therefore critical to develop techniques that can reve…

AttributecounterfactualFairnessSensitivity

Automated Ableism: An Exploration of Explicit Disability Biases in Sentiment and Toxicity Analysis Models

2023-07-18 · Pranav Narayanan Venkit, Mukund Srinath, Shomir Wilson

We analyze sentiment analysis and toxicity detection models to detect the presence of explicit bias against people with disability (PWD). We employ the bias identification framework of Perturbation Sensitivity Analysis t…

Sentiment Analysis

Cyberbullying Detection with Fairness Constraints

2020-05-09 · Oguzhan Gencoglu

Cyberbullying is a widespread adverse phenomenon among online social interactions in today's digital society. While numerous computational studies focus on enhancing the cyberbullying detection performance of machine lea…

BIG-bench Machine LearningFairness

Detecting Unintended Social Bias in Toxic Language Datasets

2022-10-21 · Nihar Sahoo, Himanshu Gupta, Pushpak Bhattacharyya

With the rise of online hate speech, automatic detection of Hate Speech, Offensive texts as a natural language processing task is getting popular. However, very little research has been done to detect unintended social b…

A Study of Implicit Bias in Pretrained Language Models against People with Disabilities

2022-10-01 · COLING 2022 10 · Pranav Narayanan Venkit, Mukund Srinath, Shomir Wilson

Pretrained language models (PLMs) have been shown to exhibit sociodemographic biases, such as against gender and race, raising concerns of downstream biases in language technologies. However, PLMs’ biases against people …

Sensitivity