paper-with-me

홈 › Papers

Detecting Unintended Social Bias in Toxic Language Datasets

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Hate speech and offensive texts are examples of damaging online content that target or promote hatred towards a group or individual member based on their actual or perceived features of identification, such as race, religion, or sexual orientation. Sharing violent and offensive content has had a significant negative impact on society. These hate speech and offensive content generally contains societal biases in them. With the rise of online hate speech, automatic detection of such biases as a natural language processing task is getting popular. However, not much research has been done to detect unintended social bias from toxic language datasets. In this paper, we introduce a new dataset from an existing toxic language dataset, to detect social biases along with their categories and targeted groups. We then report baseline performances of both classification and generation tasks on our curated dataset using transformer-based models. Our study motivates a systematic extraction of social bias data from toxic language data.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting Unintended Social Bias in Toxic Language Datasets

2022-10-21 · Nihar Sahoo, Himanshu Gupta, Pushpak Bhattacharyya

With the rise of online hate speech, automatic detection of Hate Speech, Offensive texts as a natural language processing task is getting popular. However, very little research has been done to detect unintended social b…

Reward Modeling for Mitigating Toxicity in Transformer-based Language Models

2022-02-19 · Farshid Faal, Ketra Schmitt, Jia Yuan Yu

Transformer-based language models are able to generate fluent text and be efficiently adapted across various natural language generation tasks. However, language models that are pretrained on large unlabeled web text cor…

Language ModelingLanguage ModellingText Generation

Reducing Unintended Identity Bias in Russian Hate Speech Detection

2020-10-22 · EMNLP (ALW) 2020 11 · Nadezhda Zueva, Madina Kabirova, Pavel Kalaidin

Toxicity has become a grave problem for many online communities and has been growing across many languages, including Russian. Hate speech creates an environment of intimidation, discrimination, and may even incite some …

Hate Speech Detection

Perturbation Sensitivity Analysis to Detect Unintended Model Biases

2019-10-09 · IJCNLP 2019 11 · Vinodkumar Prabhakaran, Ben Hutchinson, Margaret Mitchell

Data-driven statistical Natural Language Processing (NLP) techniques leverage large amounts of language data to build models that can understand language. However, most language data reflect the public discourse at the t…

modelSensitivitySentiment Analysis

Determination of toxic comments and unintended model bias minimization using Deep learning approach

2023-11-08 · Md Azim Khan

Online conversations can be toxic and subjected to threats, abuse, or harassment. To identify toxic text comments, several deep learning and machine learning models have been proposed throughout the years. However, recen…

regression