paper-with-me

홈 › Papers

Detecting Unintended Social Bias in Toxic Language Datasets

2022-10-21 · Nihar Sahoo, Himanshu Gupta, Pushpak Bhattacharyya

With the rise of online hate speech, automatic detection of Hate Speech, Offensive texts as a natural language processing task is getting popular. However, very little research has been done to detect unintended social bias from these toxic language datasets. This paper introduces a new dataset ToxicBias curated from the existing dataset of Kaggle competition named "Jigsaw Unintended Bias in Toxicity Classification". We aim to detect social biases, their categories, and targeted groups. The dataset contains instances annotated for five different bias categories, viz., gender, race/ethnicity, religion, political, and LGBTQ. We train transformer-based models using our curated datasets and report baseline performance for bias identification, target generation, and bias implications. Model biases and their mitigation are also discussed in detail. Our study motivates a systematic extraction of social bias data from toxic language datasets. All the codes and dataset used for experiments in this work are publicly available

📄 PDF Abstract BibTeX arXiv:2210.11762

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting Unintended Social Bias in Toxic Language Datasets

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Hate speech and offensive texts are examples of damaging online content that target or promote hatred towards a group or individual member based on their actual or perceived features of identification, such as race, reli…

Reward Modeling for Mitigating Toxicity in Transformer-based Language Models

2022-02-19 · Farshid Faal, Ketra Schmitt, Jia Yuan Yu

Transformer-based language models are able to generate fluent text and be efficiently adapted across various natural language generation tasks. However, language models that are pretrained on large unlabeled web text cor…

Language ModelingLanguage ModellingText Generation

Reducing Unintended Identity Bias in Russian Hate Speech Detection

2020-10-22 · EMNLP (ALW) 2020 11 · Nadezhda Zueva, Madina Kabirova, Pavel Kalaidin

Toxicity has become a grave problem for many online communities and has been growing across many languages, including Russian. Hate speech creates an environment of intimidation, discrimination, and may even incite some …

Hate Speech Detection

Perturbation Sensitivity Analysis to Detect Unintended Model Biases

2019-10-09 · IJCNLP 2019 11 · Vinodkumar Prabhakaran, Ben Hutchinson, Margaret Mitchell

Data-driven statistical Natural Language Processing (NLP) techniques leverage large amounts of language data to build models that can understand language. However, most language data reflect the public discourse at the t…

modelSensitivitySentiment Analysis

Determination of toxic comments and unintended model bias minimization using Deep learning approach

2023-11-08 · Md Azim Khan

Online conversations can be toxic and subjected to threats, abuse, or harassment. To identify toxic text comments, several deep learning and machine learning models have been proposed throughout the years. However, recen…

regression