paper-with-me

홈 › Papers

On Gender Biases in Offensive Language Classification Models

2022-07-01 · NAACL (GeBNLP) 2022 7 · Sanjana Marcé, Adam Poliak

We explore whether neural Natural Language Processing models trained to identify offensive language in tweets contain gender biases. We add historically gendered and gender ambiguous American names to an existing offensive language evaluation set to determine whether models? predictions are sensitive or robust to gendered names. While we see some evidence that these models might be prone to biased stereotypes that men use more offensive language than women, our results indicate that these models? binary predictions might not greatly change based upon gendered names.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Classification

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Towards Equal Gender Representation in the Annotations of Toxic Language Detection

2021-06-04 · ACL (GeBNLP) 2021 8 · Elizabeth Excell, Noura Al Moubayed

Classifiers tend to propagate biases present in the data on which they are trained. Hence, it is important to understand how the demographic identities of the annotators of comments affect the fairness of the resulting m…

Fairness

Towards Safer AI Moderation: Evaluating LLM Moderators Through a Unified Benchmark Dataset and Advocating a Human-First Approach

2025-08-09 · Naseem Machlovi, Maryam Saleki, Innocent Ababio, Ruhul Amin arxiv

As AI systems become more integrated into daily life, the need for safer and more reliable moderation has never been greater. Large Language Models (LLMs) have demonstrated remarkable capabilities, surpassing earlier mod…

Detecting Unintended Social Bias in Toxic Language Datasets

2022-10-21 · Nihar Sahoo, Himanshu Gupta, Pushpak Bhattacharyya

With the rise of online hate speech, automatic detection of Hate Speech, Offensive texts as a natural language processing task is getting popular. However, very little research has been done to detect unintended social b…

Should We Attend More or Less? Modulating Attention for Fairness

2023-05-22 · Abdelrahman Zayed, Goncalo Mordido, Samira Shabanian, Sarath Chandar

The advances in natural language processing (NLP) pose both opportunities and challenges. While recent progress enables the development of high-performing models for a variety of tasks, it also poses the risk of models l…

Fairnesstext-classificationText Classification

Detecting Unintended Social Bias in Toxic Language Datasets

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Hate speech and offensive texts are examples of damaging online content that target or promote hatred towards a group or individual member based on their actual or perceived features of identification, such as race, reli…