paper-with-me

Papers

Revealing Temporal Label Noise in Multimodal Hateful Video Classification

2025-08-06 · Shuonan Yang, Tailin Chen, Rahul Singh, Jiangbei Yue, Jianbo Jiao, Zeyu Fu arxiv

The rapid proliferation of online multimedia content has intensified the spread of hate speech, presenting critical societal and regulatory challenges. While recent work has advanced multimodal hateful video detection, most approaches rely on coarse, video-level annotations that overlook the temporal granularity of hateful content. This introduces substantial label noise, as videos annotated as hateful often contain long non-hateful segments. In this paper, we investigate the impact of such label ambiguity through a fine-grained approach. Specifically, we trim hateful videos from the HateMM and MultiHateClip English datasets using annotated timestamps to isolate explicitly hateful segments. We then conduct an exploratory analysis of these trimmed segments to examine the distribution and characteristics of both hateful and non-hateful content. This analysis highlights the degree of semantic overlap and the confusion introduced by coarse, video-level annotations. Finally, controlled experiments demonstrated that time-stamp noise fundamentally alters model decision boundaries and weakens classification confidence, highlighting the inherent context dependency and temporal continuity of hate speech expression. Our findings provide new insights into the temporal dynamics of multimodal hateful videos and highlight the need for temporally aware models and benchmarks for improved robustness and interpretability. Code and data are available at https://github.com/Multimodal-Intelligence-Lab-MIL/HatefulVideoLabelNoise.

📄 PDF Abstract BibTeX arXiv:2508.04900

Code (0)

등록된 구현이 없습니다.

Tasks

Video Classification

Similar Papers 제목 키워드 기반

HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection

2025-08-03 · Han Wang, Zhuoran Wang, Roy Ka-Wei Lee arxiv

Detecting hate speech in videos remains challenging due to the complexity of multimodal content and the lack of fine-grained annotations in existing datasets. We present HateClipSeg, a large-scale multimodal dataset with…

Video Classification

CLARA: Clip-Level Multimodal Alignment with VLM-Derived Rationales for Hateful Video Detection

2026-08-16 · Yuchen Zhang, Shuang Dai, Zeyu Fu, Yunfei Long 외 arxiv

Hateful video detection has become increasingly important with the rapid growth of video-centric social media platforms, given the serious risks that hate speech poses to both individual well-being and social cohesion. C…

Multimodal Hate Detection Using Dual-Stream Graph Neural Networks

2025-09-16 · Jiangbei Yue, Shuonan Yang, Tailin Chen, Jianbo Jiao 외 arxiv

Hateful videos present serious risks to online safety and real-world well-being, necessitating effective detection methods. Although multimodal classification approaches integrating information from several modalities ou…

Video ClassificationGraph Neural Network

Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation

2026-03-24 · Nils A. Herrmann, Tobias Eder, Jingyi He, Georg Groh arxiv

Current multimodal toxicity benchmarks typically use a single binary hatefulness label. This coarse approach conflates two fundamentally different characteristics of expression: tone and content. Drawing on communication…

Transfer Learning

MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili

2024-07-28 · Han Wang, Tan Rui Yang, Usman Naseem, Roy Ka-Wei Lee

Hate speech is a pressing issue in modern society, with significant effects both online and offline. Recent research in hate speech detection has primarily centered on text-based media, largely overlooking multimodal con…

Hate Speech DetectionVideo Classification