paper-with-me

Papers

Watching the Watchers: A Comparative Fairness Audit of Cloud-based Content Moderation Services

2024-06-20 · David Hartmann, Amin Oueslati, Dimitri Staufer

Online platforms face the challenge of moderating an ever-increasing volume of content, including harmful hate speech. In the absence of clear legal definitions and a lack of transparency regarding the role of algorithms in shaping decisions on content moderation, there is a critical need for external accountability. Our study contributes to filling this gap by systematically evaluating four leading cloud-based content moderation services through a third-party audit, highlighting issues such as biases against minorities and vulnerable groups that may arise through over-reliance on these services. Using a black-box audit approach and four benchmark data sets, we measure performance in explicit and implicit hate speech detection as well as counterfactual fairness through perturbation sensitivity analysis and present disparities in performance for certain target identity groups and data sets. Our analysis reveals that all services had difficulties detecting implicit hate speech, which relies on more subtle and codified messages. Moreover, our results point to the need to remove group-specific bias. It seems that biases towards some groups, such as Women, have been mostly rectified, while biases towards other groups, such as LGBTQ+ and PoC remain.

📄 PDF Abstract BibTeX arXiv:2406.14154

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualFairnessHate Speech Detection

Similar Papers 제목 키워드 기반

Non-Comparative Fairness for Human-Auditing and Its Relation to Traditional Fairness Notions

2021-06-29 · Mukund Telukunta, Venkata Sriram Siddhardh Nadendla

Bias evaluation in machine-learning based services (MLS) based on traditional algorithmic fairness notions that rely on comparative principles is practically difficult, making it necessary to rely on human auditor feedba…

FairnessRelation

On the Identification of Fair Auditors to Evaluate Recommender Systems based on a Novel Non-Comparative Fairness Notion

2020-09-09 · Mukund Telukunta, Venkata Sriram Siddhardh Nadendla

Decision-support systems are information systems that offer support to people's decisions in various applications such as judiciary, real-estate and banking sectors. Lately, these support systems have been found to be di…

FairnessRecommendation Systems

Fairness Score and Process Standardization: Framework for Fairness Certification in Artificial Intelligence Systems

2022-01-10 · Avinash Agarwal, Harsh Agarwal, Nihaarika Agarwal

Decisions made by various Artificial Intelligence (AI) systems greatly influence our day-to-day lives. With the increasing use of AI systems, it becomes crucial to know that they are fair, identify the underlying biases …

Decision MakingFairness

Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers

2025-01-23 · Akshit Achara, Anshuman Chhabra

AI Safety Moderation (ASM) classifiers are designed to moderate content on social media platforms and to serve as guardrails that prevent Large Language Models (LLMs) from being fine-tuned on unsafe inputs. Owing to thei…

Fairness

Auditing and Achieving Intersectional Fairness in Classification Problems

2019-11-04 · Giulio Morina, Viktoriia Oliinyk, Julian Waton, Ines Marusic 외

Machine learning algorithms are extensively used to make increasingly more consequential decisions about people, so achieving optimal predictive performance can no longer be the only focus. A particularly important consi…

AttributeBIG-bench Machine LearningClassificationFairness+1