paper-with-me

Papers

Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos

2026-06-10 · Junyu Lu, Deyi Ji, Liqun Liu, Xiaokun Zhang, Youlin Wu, Roy Ka-Wei Lee, Peng Shu, Huan Yu, Jie Jiang, Bo Xu, Liang Yang, Hongfei Lin arxiv

Hateful videos have become prevalent on online platforms, highlighting an urgent need for effective detection. However, existing studies primarily focus on binary classification and fail to provide contextual rationales that reveal the implicit meanings behind these judgments, significantly undermining model explainability. To fill this gap, we aim to achieve explainable hateful video detection, enabling models to provide contextual rationales that integrate relevant evidence and logical reasoning alongside decisions. This approach can comprehensively enhance the understanding of video content and the explainability of the decision-making process. We first introduce two datasets, Ex-HateMM and Ex-ImpliHateVid, for explainable hateful video detection. Each dataset provides fine-grained annotations of multimodal harmful elements, along with contextual rationales. We then propose an Information Augmentation and Reasoning Enhancement (IARE) framework designed for explainable detection. The framework employs an information augmentation phase that leverages the multimodal chain-of-thought to integrate harmful elements, thereby enriching rationale evidence. Additionally, IARE incorporates a reasoning enhancement phase, in which Direct Preference Optimization guides the model toward correct reasoning paths and away from incorrect ones, thereby improving the logical coherence of its justifications. We conduct extensive experiments on the two datasets, comparing multiple baselines with our proposed IARE framework. The results demonstrate that IARE achieves state-of-the-art performance while also generating accurate rationales.

📄 PDF Abstract BibTeX arXiv:2606.11953

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationLogical Reasoning

Similar Papers 제목 키워드 기반

Grounding Task Assistance with Multimodal Cues from a Single Demonstration

2025-05-02 · Gabriel Sarch, Balasaravanan Thoravi Kumaravel, Sahithya Ravi, Vibhav Vineet 외

A person's demonstration often serves as a key reference for others learning the same task. However, RGB video, the dominant medium for representing these demonstrations, often fails to capture fine-grained contextual cu…

Question AnsweringVisual Question Answering

More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection

2026-03-22 · Runze Sun, Yu Zheng, Zexuan Xiong, Zhongjin Qu 외 arxiv

Combating hate speech on social media is critical for securing cyberspace, yet relies heavily on the efficacy of automated detection systems. As content formats evolve, hate speech is transitioning from solely plain text…

Hate Speech DetectionBinary Classification

Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text

2026-02-04 · Ahmed Ruby, Christian Hardmeier, Sara Stymne arxiv

Implicit discourse relation classification is a challenging task, as it requires inferring meaning from context. While contextual cues can be distributed across modalities and vary across languages, they are not always c…

Relation ClassificationCross-Lingual Transfer

CIEC: Coupling Implicit and Explicit Cues for Multimodal Weakly Supervised Manipulation Localization

2026-02-02 · Xinquan Yu, Wei Lu, Xiangyang Luo, Rui Yang arxiv

To mitigate the threat of misinformation, multimodal manipulation localization has garnered growing attention. Consider that current methods rely on costly and time-consuming fine-grained annotations, such as patch/token…

Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language Models

2023-12-09 · Hongzhan Lin, Ziyang Luo, Jing Ma, Long Chen

The age of social media is rife with memes. Understanding and detecting harmful memes pose a significant challenge due to their implicit meaning that is not explicitly conveyed through the surface text and image. However…

Multimodal Reasoning