paper-with-me

Papers

Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models

2024-01-24 · Hongzhan Lin, Ziyang Luo, Wei Gao, Jing Ma, Bo wang, Ruichao Yang

The age of social media is flooded with Internet memes, necessitating a clear grasp and effective identification of harmful ones. This task presents a significant challenge due to the implicit meaning embedded in memes, which is not explicitly conveyed through the surface text and image. However, existing harmful meme detection methods do not present readable explanations that unveil such implicit meaning to support their detection decisions. In this paper, we propose an explainable approach to detect harmful memes, achieved through reasoning over conflicting rationales from both harmless and harmful positions. Specifically, inspired by the powerful capacity of Large Language Models (LLMs) on text generation and reasoning, we first elicit multimodal debate between LLMs to generate the explanations derived from the contradictory arguments. Then we propose to fine-tune a small language model as the debate judge for harmfulness inference, to facilitate multimodal fusion between the harmfulness rationales and the intrinsic multimodal information within memes. In this way, our model is empowered to perform dialectical reasoning over intricate and implicit harm-indicative patterns, utilizing multimodal explanations originating from both harmless and harmful arguments. Extensive experiments on three public meme datasets demonstrate that our harmful meme detection approach achieves much better performance than state-of-the-art methods and exhibits a superior capacity for explaining the meme harmfulness of the model predictions.

📄 PDF Abstract BibTeX arXiv:2401.13298

Code (1)

hkbunlp/explainhm-www2024 공식 구현 pytorch

Tasks

Hateful Meme ClassificationLanguage ModellingSmall Language ModelText Generation

Similar Papers 제목 키워드 기반

Multimodal and Explainable Internet Meme Classification

2022-12-11 · Abhinav Kumar Thakur, Filip Ilievski, Hông-Ân Sandlin, Zhivar Sourati 외

In the current context where online platforms have been effectively weaponized in a variety of geo-political events and social issues, Internet memes make fair content moderation at scale even more difficult. Existing wo…

ClassificationExplainable ModelsHate Speech DetectionMeme Classification

Detecting Harmful Memes with Decoupled Understanding and Guided CoT Reasoning

2025-06-10 · Fengjun Pan, Anh Tuan Luu, Xiaobao Wu

Detecting harmful memes is essential for maintaining the integrity of online environments. However, current approaches often struggle with resource efficiency, flexibility, or explainability, limiting their practical dep…

Meme Classification

Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language Models

2023-12-09 · Hongzhan Lin, Ziyang Luo, Jing Ma, Long Chen

The age of social media is rife with memes. Understanding and detecting harmful memes pose a significant challenge due to their implicit meaning that is not explicitly conveyed through the surface text and image. However…

Multimodal Reasoning

Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

2026-07-16 · Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani 외 arxiv

Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where humor, sarcasm, and harmful intent coexist. These complexities highlight …

Humor Detection

Multimodal Learning for Hateful Memes Detection

2020-11-25 · Yi Zhou, Zhenhao Chen

Memes are used for spreading ideas through social networks. Although most memes are created for humor, some memes become hateful under the combination of pictures and text. Automatically detecting the hateful memes can h…

Image CaptioningMultimodal Deep Learning