paper-with-me

Papers

Are you sure? Measuring models bias in content moderation through uncertainty

2025-09-21 · Alessandra Urbinati, Mirko Lai, Simona Frenda, Marco Antonio Stranisci arxiv

Automatic content moderation is crucial to ensuring safety in social media. Language Model-based classifiers are being increasingly adopted for this task, but it has been shown that they perpetuate racial and social biases. Even if several resources and benchmark corpora have been developed to challenge this issue, measuring the fairness of models in content moderation remains an open issue. In this work, we present an unsupervised approach that benchmarks models on the basis of their uncertainty in classifying messages annotated by people belonging to vulnerable groups. We use uncertainty, computed by means of the conformal prediction technique, as a proxy to analyze the bias of 11 models against women and non-white annotators and observe to what extent it diverges from metrics based on performance, such as the $F_1$ score. The results show that some pre-trained models predict with high accuracy the labels coming from minority groups, even if the confidence in their prediction is low. Therefore, by measuring the confidence of models, we are able to see which groups of annotators are better represented in pre-trained models and lead the debiasing process of these models before their effective use.

📄 PDF Abstract BibTeX arXiv:2509.22699

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services

2024-06-20 · David Hartmann, Amin Oueslati, Dimitri Staufer

Online platforms face the challenge of moderating an ever-increasing volume of content, including harmful hate speech. In the absence of clear legal definitions and a lack of transparency regarding the role of algorithms…

counterfactualFairnessHate Speech Detection

Like trainer, like bot? Inheritance of bias in algorithmic content moderation

2017-07-05 · Reuben Binns, Michael Veale, Max Van Kleek, Nigel Shadbolt

The internet has become a central medium through which `networked publics' express their opinions and engage in debate. Offensive comments and personal attacks can inhibit participation in these spaces. Automated content…

Navigate

BiasX: "Thinking Slow" in Toxic Content Moderation with Explanations of Implied Social Biases

2023-05-23 · Yiming Zhang, Sravani Nanduri, Liwei Jiang, Tongshuang Wu 외

Toxicity annotators and content moderators often default to mental shortcuts when making decisions. This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected. We introduc…

Gender and content bias in Large Language Models: a case study on Google Gemini 2.0 Flash Experimental

2025-03-18 · Roberto Balestri

This study evaluates the biases in Gemini 2.0 Flash Experimental, a state-of-the-art large language model (LLM) developed by Google, focusing on content moderation and gender disparities. By comparing its performance to …

FairnessLarge Language Model

Bandits for Online Calibration: An Application to Content Moderation on Social Media Platforms

2022-11-11 · Vashist Avadhanula, Omar Abdul Baki, Hamsa Bastani, Osbert Bastani 외

We describe the current content moderation strategy employed by Meta to remove policy-violating content from its platforms. Meta relies on both handcrafted and learned risk models to flag potentially violating content fo…