paper-with-me

홈 › Papers

Revisiting the "Video" in Video-Language Understanding

2022-06-03 · CVPR 2022 1 · Shyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, Juan Carlos Niebles

What makes a video task uniquely suited for videos, beyond what can be understood from a single image? Building on recent progress in self-supervised image-language models, we revisit this question in the context of video and language tasks. We propose the atemporal probe (ATP), a new model for video-language analysis which provides a stronger bound on the baseline accuracy of multimodal models constrained by image-level understanding. By applying this model to standard discriminative video and language tasks, such as video question answering and text-to-video retrieval, we characterize the limitations and potential of current video-language benchmarks. We find that understanding of event temporality is often not necessary to achieve strong or state-of-the-art performance, even compared with recent large-scale video-language models and in contexts intended to benchmark deeper video-level understanding. We also demonstrate how ATP can improve both video-language dataset and model design. We describe a technique for leveraging ATP to better disentangle dataset subsets with a higher concentration of temporally challenging data, improving benchmarking efficacy for causal and temporal understanding. Further, we show that effectively integrating ATP into full video-level temporal models can improve efficiency and state-of-the-art accuracy.

📄 PDF Abstract BibTeX arXiv:2206.01720

Code (1)

stanfordvl/atp-video-language pytorch

Tasks

BenchmarkingQuestion AnsweringRetrievalText to Video RetrievalVideo Question AnsweringVideo Retrieval

Similar Papers 제목 키워드 기반

SPOT! Revisiting Video-Language Models for Event Understanding

2023-11-21 · Gengyuan Zhang, Jinhe Bi, Jindong Gu, Yanyu Chen 외

Understanding videos is an important research topic for multimodal learning. Leveraging large-scale datasets of web-crawled video-text pairs as weak supervision has become a pre-training paradigm for learning joint repre…

AttributeVideo Understanding

Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models

2026-03-18 · Linghao Zhang, Jungang Li, Yonghua Hei, Sicheng Tao 외 arxiv

Multimodal large language models (MLLMs) are typically trained in multiple stages, with video-based supervised fine-tuning (Video-SFT) serving as a key step for improving visual understanding. Yet its effect on the fine-…

Revisiting Kernel Temporal Segmentation as an Adaptive Tokenizer for Long-form Video Understanding

2023-09-20 · Mohamed Afham, Satya Narayan Shukla, Omid Poursaeed, Pengchuan Zhang 외

While most modern video understanding models operate on short-range clips, real-world videos are often several minutes long with semantically consistent segments of variable length. A common approach to process long vide…

Action LocalizationFormTemporal Action LocalizationVideo Classification+1

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

2026-06-23 · Yixuan Li, Guangzhi Sun, Yudong Yang, Chao Zhang arxiv

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question …

Reinforcement LearningQuestion Answering

Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding

2025-08-28 · Gowreesh Mago, Pascal Mettes, Stevan Rudinac arxiv

The automatic understanding of video content is advancing rapidly. Empowered by deeper neural networks and large datasets, machines are increasingly capable of understanding what is concretely visible in video frames, wh…