StopHC: A Harmful Content Detection and Mitigation Architecture for Social Media Platforms
The mental health of social media users has started more and more to be put at risk by harmful, hateful, and offensive content. In this paper, we propose \textsc{StopHC}, a harmful content detection and mitigation architecture for social media platforms. Our aim with \textsc{StopHC} is to create more secure online environments. Our solution contains two modules, one that employs deep neural network architecture for harmful content detection, and one that uses a network immunization algorithm to block toxic nodes and stop the spread of harmful content. The efficacy of our solution is demonstrated by experiments conducted on two real-world datasets.
Code (1)
Similar Papers 제목 키워드 기반
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
Large language models (LLMs) have become ubiquitous, thus it is important to understand their risks and limitations. Smaller LLMs can be deployed where compute resources are constrained, such as edge devices, but with di…
HOD: A Benchmark Dataset for Harmful Object Detection
Recent multi-media data such as images and videos have been rapidly spread out on various online services such as social network services (SNS). With the explosive growth of online media services, the number of image con…
Objectobject-detectionObject DetectionHarmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
Short video platforms, such as YouTube, Instagram, or TikTok, are used by billions of users globally. These platforms expose users to harmful content, ranging from clickbait or physical harms to misinformation or online …
Binary ClassificationMisinformationtext annotationFrom Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation
Large language models (LLMs) demonstrate impressive performance across a wide range of general NLP tasks; however, their effectiveness in sensitive domains, such as hate speech detection, remains less clear. Prior studie…
Hate Speech DetectionNDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation
Despite the impressive generative capabilities of text-to-image (T2I) diffusion models, they remain vulnerable to generating inappropriate content, especially when confronted with implicit sexual prompts. Unlike explicit…
Text-to-Image Generation