paper-with-me

Papers

Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework

2026-05-02 · Ewelina Gajewska, Michal Wawer, Katarzyna Budzynska, Jaroslaw A. Chudziak arxiv

The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy. Traditional content moderation systems rely on centralised, top-down rules, often failing to accommodate the subjective nature of harm perception. This paper proposes an LLM-based multi-agent personalised inference framework that filters content based on unique sensitivity profiles of individual users. Our architecture combines domain-specific Expert Agents, a Manager Agent for orchestrating content analysis and agent selection, and a Ghost Profile Agent for simulating user perspectives, to inform moderation decisions. Evaluated against a range of non-personalised baselines, the system demonstrates up to a 32% improvement in accuracy, showing increased alignment with individual user sensitivities. Beyond technical performance, our framework provides policy-relevant insights for platform governance, providing a scalable way to reconcile moderation policies with societal and individual digital rights

📄 PDF Abstract BibTeX arXiv:2605.01416

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting Harmful Content On Online Platforms: What Platforms Need Vs. Where Research Efforts Go

2021-02-27 · Arnav Arora, Preslav Nakov, Momchil Hardalov, Sheikh Muhammad Sarwar 외

The proliferation of harmful content on online platforms is a major societal problem, which comes in many different forms including hate speech, offensive language, bullying and harassment, misinformation, spam, violence…

Abusive LanguageMisinformation

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks

2026-06-08 · Qin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang 외 arxiv

Large language model (LLM)-powered content moderation systems are a critical defense against harmful online content. However, they operate primarily on tokenized text and often overlook visual cues that humans naturally …

A Unified Taxonomy of Harmful Content

2020-11-01 · EMNLP (ALW) 2020 11 · Michele Banko, Brendon MacKeen, Laurie Ray

The ability to recognize harmful content within online communities has come into focus for researchers, engineers and policy makers seeking to protect users from abuse. While the number of datasets aiming to capture form…

Reliable Decision from Multiple Subtasks through Threshold Optimization: Content Moderation in the Wild

2022-08-16 · Donghyun Son, Byounggyu Lew, Kwanghee Choi, Yongsu Baek 외

Social media platforms struggle to protect users from harmful content through content moderation. These platforms have recently leveraged machine learning models to cope with the vast amount of user-generated content dai…

Metamorphic Testing for Audio Content Moderation Software

2025-09-29 · Wenxuan Wang, Yongjiang Wu, Junyuan Zhang, Shuqing Li 외 arxiv

The rapid growth of audio-centric platforms and applications such as WhatsApp and Twitter has transformed the way people communicate and share audio content in modern society. However, these platforms are increasingly mi…