paper-with-me

Papers

StopHC: A Harmful Content Detection and Mitigation Architecture for Social Media Platforms

2024-11-09 · Ciprian-Octavian Truică, Ana-Teodora Constantinescu, Elena-Simona Apostol

The mental health of social media users has started more and more to be put at risk by harmful, hateful, and offensive content. In this paper, we propose \textsc{StopHC}, a harmful content detection and mitigation architecture for social media platforms. Our aim with \textsc{StopHC} is to create more secure online environments. Our solution contains two modules, one that employs deep neural network architecture for harmful content detection, and one that uses a network immunization algorithm to block toxic nodes and stop the spread of harmful content. The efficacy of our solution is demonstrated by experiments conducted on two real-world datasets.

📄 PDF Abstract BibTeX arXiv:2411.06138

Code (1)

DS4AI-UPB/StopHC_HarmfulContentMitigation 공식 구현 tf

Similar Papers 제목 키워드 기반

Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet

2025-02-07 · Berk Atıl, Vipul Gupta, Sarkar Snigdha Sarathi Das, Rebecca J. Passonneau

Large language models (LLMs) have become ubiquitous, thus it is important to understand their risks and limitations. Smaller LLMs can be deployed where compute resources are constrained, such as edge devices, but with di…

HOD: A Benchmark Dataset for Harmful Object Detection

2023-10-08 · Eungyeom Ha, Heemook Kim, Sung Chul Hong, Dongbin Na

Recent multi-media data such as images and videos have been rapidly spread out on various online services such as social network services (SNS). With the explosive growth of online media services, the number of image con…

Objectobject-detectionObject Detection

Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators

2024-11-06 · Claire Wonjeong Jo, Miki Wesołowska, Magdalena Wojcieszak

Short video platforms, such as YouTube, Instagram, or TikTok, are used by billions of users globally. These platforms expose users to harmful content, ranging from clickbait or physical harms to misinformation or online …

Binary ClassificationMisinformationtext annotation

From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation

2026-08-26 · Lukas Edman, Daryna Dementieva, Alexander Fraser arxiv

Large language models (LLMs) demonstrate impressive performance across a wide range of general NLP tasks; however, their effectiveness in sensitive domains, such as hate speech detection, remains less clear. Prior studie…

Hate Speech Detection

NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation

2025-10-17 · Yitong Sun, Yao Huang, Ruochen Zhang, Huanran Chen 외 arxiv

Despite the impressive generative capabilities of text-to-image (T2I) diffusion models, they remain vulnerable to generating inappropriate content, especially when confronted with implicit sexual prompts. Unlike explicit…

Text-to-Image Generation