Enabling Contextual Soft Moderation on Social Media through Contrastive Textual Deviation
Automated soft moderation systems are unable to ascertain if a post supports or refutes a false claim, resulting in a large number of contextual false positives. This limits their effectiveness, for example undermining trust in health experts by adding warnings to their posts or resorting to vague warnings instead of granular fact-checks, which result in desensitizing users. In this paper, we propose to incorporate stance detection into existing automated soft-moderation pipelines, with the goal of ruling out contextual false positives and providing more precise recommendations for social media content that should receive warnings. We develop a textual deviation task called Contrastive Textual Deviation (CTD) and show that it outperforms existing stance detection approaches when applied to soft moderation.We then integrate CTD into the stateof-the-art system for automated soft moderation Lambretta, showing that our approach can reduce contextual false positives from 20% to 2.1%, providing another important building block towards deploying reliable automated soft moderation tools on social media.
Code (0)
등록된 구현이 없습니다.
Tasks
Stance DetectionSimilar Papers 제목 키워드 기반
Validating Multimedia Content Moderation Software via Semantic Fusion
The exponential growth of social media platforms, such as Facebook and TikTok, has revolutionized communication and content publication in human society. Users on these platforms can publish multimedia content that deliv…
Sentencesoftware testingAn Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation Software
The exponential growth of social media platforms has brought about a revolution in communication and content dissemination in human society. Nevertheless, these platforms are being increasingly misused to spread toxic co…
MTTM: Metamorphic Testing for Textual Content Moderation Software
The exponential growth of social media platforms such as Twitter and Facebook has revolutionized textual communication and textual content publication in human society. However, they have been increasingly exploited to p…
SentenceTAR on Social Media: A Framework for Online Content Moderation
Content moderation (removing or limiting the distribution of posts based on their contents) is one tool social networks use to fight problems such as harassment and disinformation. Manually screening all content is usual…
Active LearningRetrievalTARBandits for Online Calibration: An Application to Content Moderation on Social Media Platforms
We describe the current content moderation strategy employed by Meta to remove policy-violating content from its platforms. Meta relies on both handcrafted and learned risk models to flag potentially violating content fo…