paper-with-me

홈 › Papers

A Video Is Not Worth a Thousand Words

2025-10-27 · Sam Pollard, Michael Wray arxiv

As we become increasingly dependent on vision language models (VLMs) to answer questions about the world around us, there is a significant amount of research devoted to increasing both the difficulty of video question answering (VQA) datasets, and the context lengths of the models that they evaluate. The reliance on large language models as backbones has lead to concerns about potential text dominance, and the exploration of interactions between modalities is underdeveloped. How do we measure whether we're heading in the right direction, with the complexity that multi-modal models introduce? We propose a joint method of computing both feature attributions and modality scores based on Shapley values, where both the features and modalities are arbitrarily definable. Using these metrics, we compare $6$ VLM models of varying context lengths on $4$ representative datasets, focusing on multiple-choice VQA. In particular, we consider video frames and whole textual elements as equal features in the hierarchy, and the multiple-choice VQA task as an interaction between three modalities: video, question and answer. Our results demonstrate a dependence on text and show that the multiple-choice VQA task devolves into a model's ability to ignore distractors. Code available at https://github.com/sjpollard/a-video-is-not-worth-a-thousand-words.

📄 PDF Abstract BibTeX arXiv:2510.23253

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation

2024-12-24 · Faraz Waseem, Muhammad Shahzad

An image may convey a thousand words, but a video composed of hundreds or thousands of image frames tells a more intricate story. Despite significant progress in multimodal large language models (MLLMs), generating exten…

Video Generation

One Picture is Worth a Thousand Words: A New Wallet Recovery Process

2022-05-05 · Hervé Chabannne, Vincent Despiegel, Linda Guiga

We introduce a new wallet recovery process. Our solution associates 1) visual passwords: a photograph of a secretly picked object (Chabanne et al., 2013) with 2) ImageNet classifiers transforming images into binary vecto…

Retrieval

Vript: A Video Is Worth Thousands of Words

2024-06-10 · Dongjie Yang, Suyuan Huang, Chengqiang Lu, Xiaodong Han 외

Advancements in multimodal learning, particularly in video understanding and generation, require high-quality video-text datasets for improved model performance. Vript addresses this issue with a meticulously annotated c…

Video CaptioningVideo Understanding

Image and Information

2016-02-03 · Frank Nielsen

A well-known old adage says that {\em "A picture is worth a thousand words!"} (attributed to the Chinese philosopher Confucius ca 500 years BC). But more precisely, what do we mean by information in images? And how can i…

Universal Differential Equations for Scientific Machine Learning

2020-01-13 · Christopher Rackauckas, Yingbo Ma, Julius Martensen, Collin Warner 외

In the context of science, the well-known adage "a picture is worth a thousand words" might well be "a model is worth a thousand datasets." In this manuscript we introduce the SciML software ecosystem as a tool for mixin…

BIG-bench Machine LearningGPU