paper-with-me

홈 › Papers

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation

2025-07-16 · Sahid Hossain Mustakim, S M Jishanul Islam, Ummay Maria Muna, Montasir Chowdhury, Mohammed Jawwadul Islam, Sadia Ahmmed, Tashfia Sikder, Syed Tasdid Azam Dhrubo, Swakkhar Shatabda

Multimodal Large Language Models (MLLMs) are increasingly used for content moderation, yet their robustness in short-form video contexts remains underexplored. Current safety evaluations often rely on unimodal attacks, failing to address combined attack vulnerabilities. In this paper, we introduce a comprehensive framework for evaluating the tri-modal safety of MLLMs. First, we present the Short-Video Multimodal Adversarial (SVMA) dataset, comprising diverse short-form videos with human-guided synthetic adversarial attacks. Second, we propose ChimeraBreak, a novel tri-modal attack strategy that simultaneously challenges visual, auditory, and semantic reasoning pathways. Extensive experiments on state-of-the-art MLLMs reveal significant vulnerabilities with high Attack Success Rates (ASR). Our findings uncover distinct failure modes, showing model biases toward misclassifying benign or policy-violating content. We assess results using LLM-as-a-judge, demonstrating attack reasoning efficacy. Our dataset and findings provide crucial insights for developing more robust and safe MLLMs.

📄 PDF Abstract BibTeX arXiv:2507.11968

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OWL (Observe, Watch, Listen): Audiovisual Temporal Context for Localizing Actions in Egocentric Videos

2022-02-10 · Merey Ramazanova, Victor Escorcia, Fabian Caba Heilbron, Chen Zhao 외

Egocentric videos capture sequences of human activities from a first-person perspective and can provide rich multimodal signals. However, most current localization methods use third-person videos and only incorporate vis…

Action LocalizationTemporal Action LocalizationTemporal Localization

Watch, Listen, and Describe: Globally and Locally Aligned Cross-Modal Attentions for Video Captioning

2018-04-15 · NAACL 2018 6 · Xin Wang, Yuan-Fang Wang, William Yang Wang

A major challenge for video captioning is to combine audio and visual cues. Existing multi-modal fusion methods have shown encouraging results in video understanding. However, the temporal structures of multiple modaliti…

Video CaptioningVideo Understanding

Watch, Listen and Tell: Multi-modal Weakly Supervised Dense Event Captioning

2019-09-22 · ICCV 2019 10 · Tanzila Rahman, Bicheng Xu, Leonid Sigal

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning…

Sound Source Localization

Listening Deepfake Detection: A New Perspective Beyond Speaking-Centric Forgery Analysis

2026-04-14 · Miao Liu, Fangda Wei, Jing Wang, Xinyuan Qian arxiv

Existing deepfake detection research has primarily focused on scenarios where the manipulated subject is actively speaking, i.e., generating fabricated content by altering the speaker's appearance or voice. However, in r…

DeepFake Detection

End-to-end Listen, Look, Speak and Act

2025-10-19 · Siyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian 외 arxiv

Human interaction is inherently multimodal and full-duplex: we listen while watching, speak while acting, and fluidly adapt to turn-taking and interruptions. Realizing these capabilities is essential for building models …

Visual Question Answering