Robustness Analysis of Video-Language Models Against Visual and Language Perturbations
Joint visual and language modeling on large-scale datasets has recently shown good progress in multi-modal tasks when compared to single modal learning. However, robustness of these approaches against real-world perturbations has not been studied. In this work, we perform the first extensive robustness study of video-language models against various real-world perturbations. We focus on text-to-video retrieval and propose two large-scale benchmark datasets, MSRVTT-P and YouCook2-P, which utilize 90 different visual and 35 different text perturbations. The study reveals some interesting initial findings from the studied models: 1) models are generally more susceptible when only video is perturbed as opposed to when only text is perturbed, 2) models that are pre-trained are more robust than those trained from scratch, 3) models attend more to scene and objects rather than motion and action. We hope this study will serve as a benchmark and guide future research in robust video-language learning. The benchmark introduced in this study along with the code and datasets is available at https://bit.ly/3CNOly4.
Code (1)
Tasks
Language ModelingLanguage ModellingRetrievalText to Video RetrievalVideo RetrievalSimilar Papers 제목 키워드 기반
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
Human intelligence requires correctness and robustness, with the former being foundational for the latter. In video understanding, correctness ensures the accurate interpretation of visual content, and robustness maintai…
Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency
Large Multimodal Models (LMMs) have recently demonstrated impressive performance on general video comprehension benchmarks. Nevertheless, for broader applications, the robustness of their temporal analysis capability nee…
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
In Vision-Language-Actionf(VLA) models, robustness to real-world perturbations is critical for deployment. Existing methods target simple visual disturbances, overlooking the broader multi-modal perturbations that arise …
From Covert Hiding to Visual Editing: Robust Generative Video Steganography
Traditional video steganography methods are based on modifying the covert space for embedding, whereas we propose an innovative approach that embeds secret message within semantic feature for steganography during the vid…
Face SwappingImage SteganographyVideo EditingIntentQA: Intent Question Answering in Videos by Cognitive Context Reasoning
Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human actions (often termed the "dark matter" of social intelligence). To bridge …
Contrastive LearningQuestion Answering