paper-with-me

Papers

A Keyword Based Approach to Understanding the Overpenalization of Marginalized Groups by English Marginal Abuse Models on Twitter

2022-10-07 · Kyra Yee, Alice Schoenauer Sebag, Olivia Redfield, Emily Sheng, Matthias Eck, Luca Belli

Harmful content detection models tend to have higher false positive rates for content from marginalized groups. In the context of marginal abuse modeling on Twitter, such disproportionate penalization poses the risk of reduced visibility, where marginalized communities lose the opportunity to voice their opinion on the platform. Current approaches to algorithmic harm mitigation, and bias detection for NLP models are often very ad hoc and subject to human bias. We make two main contributions in this paper. First, we design a novel methodology, which provides a principled approach to detecting and measuring the severity of potential harms associated with a text-based model. Second, we apply our methodology to audit Twitter's English marginal abuse model, which is used for removing amplification eligibility of marginally abusive content. Without utilizing demographic labels or dialect classifiers, we are still able to detect and measure the severity of issues related to the over-penalization of the speech of marginalized communities, such as the use of reclaimed speech, counterspeech, and identity related terms. In order to mitigate the associated harms, we experiment with adding additional true negative examples and find that doing so provides improvements to our fairness metrics without large degradations in model performance.

📄 PDF Abstract BibTeX arXiv:2210.06351

Code (0)

등록된 구현이 없습니다.

Tasks

Bias DetectionFairness

Methods 이 논문이 사용한 방법론

HOC 설명 없음

Similar Papers 제목 키워드 기반

Out of Sight Out of Mind, Out of Sight Out of Mind: Measuring Bias in Language Models Against Overlooked Marginalized Groups in Regional Contexts

2025-04-17 · Fatma Elsafoury, David Hartmann

We know that language models (LMs) form biases and stereotypes of minorities, leading to unfair treatments of members of these groups, thanks to research mainly in the US and the broader English-speaking world. As the ne…

Detoxifying Language Models Risks Marginalizing Minority Voices

2021-04-13 · NAACL 2021 4 · Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan 외

Language models (LMs) must be both safe and equitable to be responsibly deployed in practice. With safety in mind, numerous detoxification techniques (e.g., Dathathri et al. 2020; Krause et al. 2020) have been proposed t…

Text Generation

Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models

2022-06-23 · NAACL 2022 7 · Yang Trista Cao, Anna Sotnikova, Hal Daumé III, Rachel Rudinger 외

NLP models trained on text have been shown to reproduce human stereotypes, which can magnify harms to marginalized groups when systems are deployed at scale. We adapt the Agency-Belief-Communion (ABC) stereotype model of…

Sensitivity

What about em? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns

2023-05-25 · Anne Lauscher, Debora Nozza, Archie Crowley, Ehm Miltersen 외

As 3rd-person pronoun usage shifts to include novel forms, e.g., neopronouns, we need more research on identity-inclusive NLP. Exclusion is particularly harmful in one of the most popular NLP applications, machine transl…

Machine TranslationTranslation

Homogeneity Bias as Differential Sampling Uncertainty in Language Models

2025-01-31 · Messi H. J. Lee, Soyeon Jeon

Prior research show that Large Language Models (LLMs) and Vision-Language Models (VLMs) represent marginalized groups more homogeneously than dominant groups. However, the mechanisms underlying this homogeneity bias rema…