paper-with-me

홈 › Papers

Time Blindness: Why Video-Language Models Can't See What Humans Can?

2025-05-30 · Ujjwal Upadhyay, Mukul Ranjan, Zhiqiang Shen, Mohamed Elhoseiny

Recent advances in vision-language models (VLMs) have made impressive strides in understanding spatio-temporal relationships in videos. However, when spatial information is obscured, these models struggle to capture purely temporal patterns. We introduce $\textbf{SpookyBench}$, a benchmark where information is encoded solely in temporal sequences of noise-like frames, mirroring natural phenomena from biological signaling to covert communication. Interestingly, while humans can recognize shapes, text, and patterns in these sequences with over 98% accuracy, state-of-the-art VLMs achieve 0% accuracy. This performance gap highlights a critical limitation: an over-reliance on frame-level spatial features and an inability to extract meaning from temporal cues. Furthermore, when trained in data sets with low spatial signal-to-noise ratios (SNR), temporal understanding of models degrades more rapidly than human perception, especially in tasks requiring fine-grained temporal reasoning. Overcoming this limitation will require novel architectures or training paradigms that decouple spatial dependencies from temporal processing. Our systematic analysis shows that this issue persists across model scales and architectures. We release SpookyBench to catalyze research in temporal pattern recognition and bridge the gap between human and machine video understanding. Dataset and code has been made available on our project website: https://timeblindness.github.io/.

📄 PDF Abstract BibTeX arXiv:2505.24867

Code (0)

등록된 구현이 없습니다.

Tasks

Temporal SequencesVideo Understanding

Similar Papers 제목 키워드 기반

Exploiting Change Blindness for Video Coding: Perspectives from a Less Promising User Study

2024-07-31 · Mitra Amiri, Steven Le Moan, Christian Herglotz

What the human visual system can perceive is strongly limited by the capacity of our working memory and attention. Such limitations result in the human observer's inability to perceive large-scale changes in a stimulus, …

Computational EfficiencyQuantizationSaliency PredictionVideo Compression

DeepBlindness: Fast Blindness Map Estimation and Blindness Type Classification for Outdoor Scene from Single Color Image

2019-11-02 · Jiaxiong Qiu, Xinyuan Yu, Guoqiang Yang, Shuaicheng Liu

Outdoor vision robotic systems and autonomous cars suffer from many image-quality issues, particularly haze, defocus blur, and motion blur, which we will define generically as "blindness issues". These blindness issues m…

General Classification

5G Edge Vision: Wearable Assistive Technology for People with Blindness and Low Vision

2023-11-23 · Tommy Azzino, Marco Mezzavilla, Sundeep Rangan, Yao Wang 외

In an increasingly visual world, people with blindness and low vision (pBLV) face substantial challenges in navigating their surroundings and interpreting visual information. From our previous work, VIS4ION is a smart we…

Lights out: training RL agents robust to temporary blindness

2023-12-05 · N. Ordonez, M. Tromp, P. M. Julbe, W. Böhmer

Agents trained with DQN rely on an observation at each timestep to decide what action to take next. However, in real world applications observations can change or be missing entirely. Examples of this could be a light bu…

The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility

2025-08-11 · Xiantao Zhang arxiv

Multimodal Large Language Models (MLLMs) hold immense promise as assistive technologies for the blind and visually impaired (BVI) community. However, we identify a critical failure mode that undermines their trustworthin…