paper-with-me

Papers

Bandits for Online Calibration: An Application to Content Moderation on Social Media Platforms

2022-11-11 · Vashist Avadhanula, Omar Abdul Baki, Hamsa Bastani, Osbert Bastani, Caner Gocmen, Daniel Haimovich, Darren Hwang, Dima Karamshuk, Thomas Leeper, Jiayuan Ma, Gregory Macnamara, Jake Mullett, Christopher Palow, Sung Park, Varun S Rajagopal, Kevin Schaeffer, Parikshit Shah, Deeksha Sinha, Nicolas Stier-Moses, Peng Xu

We describe the current content moderation strategy employed by Meta to remove policy-violating content from its platforms. Meta relies on both handcrafted and learned risk models to flag potentially violating content for human review. Our approach aggregates these risk models into a single ranking score, calibrating them to prioritize more reliable risk models. A key challenge is that violation trends change over time, affecting which risk models are most reliable. Our system additionally handles production challenges such as changing risk models and novel risk models. We use a contextual bandit to update the calibration in response to such trends. Our approach increases Meta's top-line metric for measuring the effectiveness of its content moderation strategy by 13%.

📄 PDF Abstract BibTeX arXiv:2211.06516

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Let Community Rules Be Reflected in Online Content Moderation

2024-08-21 · Wangjiaxuan Xin, Kanlun Wang, Zhe Fu, Lina Zhou

Content moderation is a widely used strategy to prevent the dissemination of irregular information on social media platforms. Despite extensive research on developing automated models to support decision-making in conten…

Decision Making

Calibrated Recommendations with Contextual Bandits

2025-09-05 · Diego Feijer, Himan Abdollahpouri, Sanket Gupta, Alexander Clare 외 arxiv

Spotify's Home page features a variety of content types, including music, podcasts, and audiobooks. However, historical data is heavily skewed toward music, making it challenging to deliver a balanced and personalized co…

Algorithmic Arbitrariness in Content Moderation

2024-02-26 · Juan Felipe Gomez, Caio Vieira Machado, Lucas Monteiro Paes, Flavio P. Calmon

Machine learning (ML) is widely used to moderate online content. Despite its scalability relative to human moderation, the use of ML introduces unique challenges to content moderation. One such challenge is predictive mu…

Content Moderation by LLM: From Accuracy to Legitimacy

2024-09-05 · Tao Huang

One trending application of LLM (large language model) is to use it for content moderation in online platforms. Most current studies on this application have focused on the metric of accuracy -- the extent to which LLMs …

Large Language Model

TAR on Social Media: A Framework for Online Content Moderation

2021-08-29 · Eugene Yang, David D. Lewis, Ophir Frieder

Content moderation (removing or limiting the distribution of posts based on their contents) is one tool social networks use to fight problems such as harassment and disinformation. Manually screening all content is usual…

Active LearningRetrievalTAR