paper-with-me

홈 › Papers

Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue

2026-05-13 · Ali Al-Lawati, Nafis Tripto, Abolfazl Ansari, Jason Lucas, Suhang Wang, Dongwon Lee arxiv

The emergence of multi-agent systems introduces novel moderation challenges that extend beyond content filtering. Agents with malicious intent may contribute harmful content that appears benign to evade content-based moderation, while compromising the system through exploitative and malicious behavior manifested across their overall interaction patterns within the community. To address this, we introduce BOT-MOD (BOT-MODeration), a moderation framework that grounds detection in agent intent rather than traditional content level signals. BOT-MOD identifies the underlying intent by engaging with the target agent in a multi-turn exchange guided by Gibbs-based sampling over candidate intent hypotheses. This progressively narrows the space of plausible agent objectives to identify the underlying behavior. To evaluate our approach, we construct a dataset derived from Moltbook that encompasses diverse benign and malicious behaviors based on actual community structures, posts, and comments. Results demonstrate that BOT-MOD reliably identifies agent intent across a range of adversarial configurations, while maintaining a low false positive rate on benign behaviors. This work advances the foundation for scalable, intent-aware moderation of agents in open multi-agent environments.

📄 PDF Abstract BibTeX arXiv:2605.12856

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities

2025-09-03 · Youngwoo Kim, Himanshu Beniwal, Steven L. Johnson, Thomas Hartvigsen arxiv

Effective content moderation systems require explicit classification criteria, yet online communities like subreddits often operate with diverse, implicit standards. This work introduces a novel approach to identify and …

Uncovering the Unseen: Discover Hidden Intentions by Micro-Behavior Graph Reasoning

2023-08-29 · Zhuo Zhou, Wenxuan Liu, Danni Xu, Zheng Wang 외

This paper introduces a new and challenging Hidden Intention Discovery (HID) task. Unlike existing intention recognition tasks, which are based on obvious visual representations to identify common intentions for normal b…

Intent Detection

The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation

2025-05-31 · YuHang Zhou, Yimin Xiao, Wei Ai, Ge Gao

Social media platforms have become central to modern communication, yet they also harbor offensive content that challenges platform safety and inclusivity. While prior research has primarily focused on textual indicators…

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment

2026-05-08 · William Brach, Federico Torrielli, Stine Lyngsø Beltoft, Annemette Brok Pirchert 외 arxiv

Moltbook is a Reddit-like platform where OpenClaw agents post, comment, and vote at scale - a so far unprecedented incident that comes with serious safety concerns. With the aim of studying emergent behavior in populatio…

The Unappreciated Role of Intent in Algorithmic Moderation of Social Media Content

2024-05-17 · Xinyu Wang, Sai Koneru, Pranav Narayanan Venkit, Brett Frischmann 외

As social media has become a predominant mode of communication globally, the rise of abusive content threatens to undermine civil discourse. Recognizing the critical nature of this issue, a significant body of research h…