paper-with-me

Papers

CLARA: Clip-Level Multimodal Alignment with VLM-Derived Rationales for Hateful Video Detection

2026-08-16 · Yuchen Zhang, Shuang Dai, Zeyu Fu, Yunfei Long, Ravi Shekhar, Haralambos Mouratidis arxiv

Hateful video detection has become increasingly important with the rapid growth of video-centric social media platforms, given the serious risks that hate speech poses to both individual well-being and social cohesion. Compared with text or static multimodal content, hateful video detection remains underexplored and significantly more challenging, as hateful meaning often arises from complex interactions among multimodal cues, including speech, audio, and visual content. Moreover, such signals are often brief, implicit, and temporally dependent, making them difficult to capture using conventional video-level representations. In this work, we propose CLARA, a clip-level multimodal framework for hateful video detection. Instead of treating a video as a single instance, CLARA models it as a sequence of fine-grained clips, enabling more precise capture of temporally localized hateful signals. We introduce a Mixture-of-Experts clip encoder for adaptive multimodal alignment, a local-global segment contrastive objective to jointly model short-term cues and long-range temporal dependencies, and VLM-derived rationales integrated via a gated Transformer to provide high-level semantic guidance. Extensive experiments on three hateful video datasets demonstrate that CLARA consistently outperforms state-of-the-art methods. Further ablation studies and parameter analyses validate the effectiveness of each component.

📄 PDF Abstract BibTeX arXiv:2608.15905

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How does longer temporal context enhance multimodal narrative video processing in the brain?

2026-02-07 · Prachi Jindal, Anant Khandelwal, Manish Gupta, Bapi S. Raju 외 arxiv

Understanding how humans and artificial intelligence systems process complex narrative videos is a fundamental challenge at the intersection of neuroscience and machine learning. This study investigates how the temporal …

Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability Perspective

2025-11-20 · Jiahao Li, Yang Lu, Yachao Zhang, Yong Xie 외 arxiv

Open-vocabulary semantic segmentation (OVSS) employs pixel-level vision-language alignment to associate category-related prompts with corresponding pixels. A key challenge is enhancing the multimodal dense prediction cap…

Semantic Segmentation

CAMEL-CLIP: Channel-aware Multimodal Electroencephalography-text Alignment for Generalizable Brain Foundation Models

2026-02-27 · Hanseul Choi, Jinyeong Park, Seongwon Jin, Sungho Park 외 arxiv

Electroencephalography (EEG) foundation models have shown promise for learning generalizable representations, yet they remain sensitive to channel heterogeneity, such as changes in channel composition or ordering. We pro…

Contrastive Learning

Spatially-Weighted CLIP for Street-View Geo-localization

2026-04-06 · Ting Han, Fengjiao Li, Chunsong Chen, Haoling Huang 외 arxiv

This paper proposes Spatially-Weighted CLIP (SW-CLIP), a novel framework for street-view geo-localization that explicitly incorporates spatial autocorrelation into vision-language contrastive learning. Unlike conventiona…

Representation LearningContrastive Learning

Toward Unified Multimodal Representation Learning for Autonomous Driving

2026-03-09 · Ximeng Tao, Dimitar Filev, Gaurav Pandey arxiv

Contrastive Language-Image Pre-training (CLIP) has shown impressive performance in aligning visual and textual representations. Recent studies have extended this paradigm to 3D vision to improve scene understanding for a…

Representation LearningContrastive LearningScene UnderstandingAutonomous Driving