paper-with-me

홈 › Papers

DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images

2025-09-26 · Dwip Dalal, Gautam Vashishtha, Anku Rani, Aishwarya Reganti, Parth Patwa, Mohd Sarique, Chandan Gupta, Keshav Nath, Viswanatha Reddy, Vinija Jain, Aman Chadha, Amitava Das, Amit Sheth, Asif Ekbal arxiv

The rise in harmful online content not only distorts public discourse but also poses significant challenges to maintaining a healthy digital environment. In response to this, we introduce a multimodal dataset uniquely crafted for identifying hate in digital content. Central to our methodology is the innovative application of watermarked, stability-enhanced, stable diffusion techniques combined with the Digital Attention Analysis Module (DAAM). This combination is instrumental in pinpointing the hateful elements within images, thereby generating detailed hate attention maps, which are used to blur these regions from the image, thereby removing the hateful sections of the image. We release this data set as a part of the dehate shared task. This paper also describes the details of the shared task. Furthermore, we present DeHater, a vision-language model designed for multimodal dehatification tasks. Our approach sets a new standard in AI-driven image hate detection given textual prompts, contributing to the development of more ethical AI applications in social media.

📄 PDF Abstract BibTeX arXiv:2509.21787

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM in the Loop: Creating the PARADEHATE Dataset for Hate Speech Detoxification

2025-06-02 · Shuzhou Yuan, Ercong Nie, Lukas Kouba, Ashish Yashwanth Kangen 외

Detoxification, the task of rewriting harmful language into non-toxic text, has become increasingly important amid the growing prevalence of toxic content online. However, high-quality parallel datasets for detoxificatio…

8k

KDEHatEval at SemEval-2019 Task 5: A Neural Network Model for Detecting Hate Speech in Twitter

2019-06-01 · SEMEVAL 2019 6 · Umme Aymun Siddiqua, Abu Nowshed Chy, Masaki Aono

In the age of emerging volume of microblog platforms, especially twitter, hate speech propagation is now of great concern. However, due to the brevity of tweets and informal user generated contents, detecting and analyzi…

SentenceSentence EmbeddingSentence-Embedding

Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models

2023-05-23 · Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes 외

State-of-the-art Text-to-Image models like Stable Diffusion and DALLE$\cdot$2 are revolutionizing how people generate visual content. At the same time, society has serious concerns about how adversaries can exploit such …

ARC-NLP at Multimodal Hate Speech Event Detection 2023: Multimodal Methods Boosted by Ensemble Learning, Syntactical and Entity Features

2023-07-25 · Umitcan Sahin, Izzet Emre Kucukkaya, Oguzhan Ozcelik, Cagri Toraman

Text-embedded images can serve as a means of spreading hate speech, propaganda, and extremist beliefs. Throughout the Russia-Ukraine war, both opposing factions heavily relied on text-embedded images as a vehicle for spr…

ARCDeep LearningEnsemble LearningEvent Detection+2

Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time Adaptation

2026-01-28 · Jiao Li, Jian Lang, Xikai Tang, Wenzheng Shu 외 arxiv

Hate Video Detection (HVD) is crucial for online ecosystems. Existing methods assume identical distributions between training (source) and inference (target) data. However, hateful content often evolves into irregular an…

Test-time Adaptation