paper-with-me

홈 › Papers

Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning

2025-10-29 · Rufan Zhang, Lin Zhang, Xianghang Mi arxiv

The proliferation of harmful online content--e.g., toxicity, spam, and negative sentiment--demands robust and adaptable moderation systems. However, prevailing moderation systems are centralized and task-specific, offering limited transparency and neglecting diverse user preferences--an approach ill-suited for privacy-sensitive or decentralized environments. We propose a novel framework that leverages in-context learning (ICL) with foundation models to unify the detection of toxicity, spam, and negative sentiment across binary, multi-class, and multi-label settings. Crucially, our approach enables lightweight personalization, allowing users to easily block new categories, unblock existing ones, or extend detection to semantic variations through simple prompt-based interventions--all without model retraining. Extensive experiments on public benchmarks (TextDetox, UCI SMS, SST2) and a new, annotated Mastodon dataset reveal that: (i) foundation models achieve strong cross-task generalization, often matching or surpassing task-specific fine-tuned models; (ii) effective personalization is achievable with as few as one user-provided example or definition; and (iii) augmenting prompts with label definitions or rationales significantly enhances robustness to noisy, real-world data. Our work demonstrates a definitive shift beyond one-size-fits-all moderation, establishing ICL as a practical, privacy-preserving, and highly adaptable pathway for the next generation of user-centric content safety systems. To foster reproducibility and facilitate future research, we publicly release our code on GitHub and the annotated Mastodon dataset on Hugging Face.

📄 PDF Abstract BibTeX arXiv:2511.05532

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Datasets for Navigating Sensitive Topics in Recommendation Systems

2025-09-08 · Amelia Kovacs, Jerry Chee, Kimia Kazemian, Sarah Dean arxiv

Personalized AI systems, from recommendation systems to chatbots, are a prevalent method for distributing content to users based on their learned preferences. However, there is growing concern about the adverse effects o…

Recommendation Systems

PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm

2026-01-13 · Jing-Jing Li, Joel Mire, Eve Fleisig, Valentina Pyatkin 외 arxiv

Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond…

IdentityGuard: Context-Aware Restriction and Provenance for Personalized Synthesis

2026-03-14 · Lingyun Zhang, Yu Xie, Ping Chen arxiv

The nature of personalized text-to-image models poses a unique safety challenge that generic context-blind methods are ill-equipped to handle. Such global filters create a dilemma: to prevent misuse, they are forced to d…

Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation

2026-04-18 · Huije Lee, Jisu Shin, Hoyun Song, Changgeon Ko 외 arxiv

Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale pre-training corpora. To address these issues, we propose a framewor…

Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models

2025-01-23 · Akash Bonagiri, Lucen Li, Rajvardhan Oak, Zeerak Babar 외

The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and scalable content moderation strategies. Current approaches rely on human moderato…

Few-Shot LearningIn-Context Learning