paper-with-me

홈 › Papers

Policy-as-Prompt: Rethinking Content Moderation in the Age of Large Language Models

2025-02-25 · Konstantina Palla, José Luis Redondo García, Claudia Hauff, Francesco Fabbri, Henrik Lindström, Daniel R. Taber, Andreas Damianou, Mounia Lalmas

Content moderation plays a critical role in shaping safe and inclusive online environments, balancing platform standards, user expectations, and regulatory frameworks. Traditionally, this process involves operationalising policies into guidelines, which are then used by downstream human moderators for enforcement, or to further annotate datasets for training machine learning moderation models. However, recent advancements in large language models (LLMs) are transforming this landscape. These models can now interpret policies directly as textual inputs, eliminating the need for extensive data curation. This approach offers unprecedented flexibility, as moderation can be dynamically adjusted through natural language interactions. This paradigm shift raises important questions about how policies are operationalised and the implications for content moderation practices. In this paper, we formalise the emerging policy-as-prompt framework and identify five key challenges across four domains: Technical Implementation (1. translating policy to prompts, 2. sensitivity to prompt structure and formatting), Sociotechnical (3. the risk of technological determinism in policy formation), Organisational (4. evolving roles between policy and machine learning teams), and Governance (5. model governance and accountability). Through analysing these challenges across technical, sociotechnical, organisational, and governance dimensions, we discuss potential mitigation approaches. This research provides actionable insights for practitioners and lays the groundwork for future exploration of scalable and adaptive content moderation systems in digital ecosystems.

📄 PDF Abstract BibTeX arXiv:2502.18695

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling

2025-08-05 · Anqi Li, Wenwei Jin, Jintao Tong, Pengda Qin 외 arxiv

Social platforms have revolutionized information sharing, but also accelerated the dissemination of harmful and policy-violating content. To ensure safety and compliance at scale, moderation systems must go beyond effici…

Structured Prediction

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

2025-01-07 · Lingzhi Yuan, Xiaojun Jia, Yihao Huang, Wei Dong 외

Text-to-image (T2I) models have been shown to be vulnerable to misuse, particularly in generating not-safe-for-work (NSFW) content, raising serious ethical concerns. In this work, we present PromptGuard, a novel content …

Image GenerationSafety Alignment

The Potential of Vision-Language Models for Content Moderation of Children's Videos

2023-12-06 · Syed Hammad Ahmed, Shengnan Hu, Gita Sukthankar

Natural language supervision has been shown to be effective for zero-shot learning in many computer vision tasks, such as object detection and activity recognition. However, generating informative prompts can be challeng…

Activity Recognitionobject-detectionZero-Shot Learning

Scaling Reinforcement Learning for Content Moderation with Large Language Models

2025-12-23 · Hamed Firooz, Rui Liu, Yuchen Lu, Zhenyu Hou 외 arxiv

Content moderation at scale remains one of the most pressing challenges in today's digital ecosystem, where billions of user- and AI-generated artifacts must be continuously evaluated for policy violations. Although rece…

Reinforcement Learning

Watch Your Language: Investigating Content Moderation with Large Language Models

2023-09-25 · Deepak Kumar, Yousef AbuHashem, Zakir Durumeric

Large language models (LLMs) have exploded in popularity due to their ability to perform a wide array of natural language tasks. Text-based content moderation is one LLM use case that has received recent enthusiasm, howe…