paper-with-me

홈 › Papers

Video-Oasis: Rethinking Evaluation of Video Understanding

2026-07-02 · Geuntaek Lim, Sungjune Park, Jaeyun Lee, Inwoong Lee, Taeoh Kim, Dongyoon Wee, Minho Shim, Yukyung Choi hf

The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have emerged to assess high-level reasoning, shared criteria for evaluating video understanding remain largely overlooked. Instead of introducing yet another benchmark, we take a step back to re-examine the criteria for evaluating video understanding. In this work, we introduce Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks. This audit reveals that 55\% of existing benchmark samples are solvable without visual input or temporal context. After filtering these shortcuts, the remaining video-native challenges expose a substantial capability gap: state-of-the-art models perform only marginally above random guessing. Building on these findings, we use the distilled challenges as a testbed to investigate which algorithmic design choices contribute to robust video understanding. We hope our work provides a practical foundation for constructing rigorous video benchmarks and evaluating future Video-LLMs. Code is available at https://github.com/sejong-rcv/Video-Oasis.

📄 PDF Abstract BibTeX arXiv:2603.29616

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

2026-04-18 · Zhijia Liang, Jiaming Li, Weikai Chen, Yanhao Zhang 외 arxiv

Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal is like an oasis-small, critical, and eas…

Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data

2024-07-18 · Wufei Ma, Kai Li, Zhongshi Jiang, Moustafa Meshry 외

Recent video-text foundation models have demonstrated strong performance on a wide variety of downstream video understanding tasks. Can these video-text models genuinely understand the contents of natural videos? Standar…

Language ModellingLarge Language ModelRetrievalVideo Understanding

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

2025-11-17 · Jiaze Li, Hao Yin, Wenhui Tan, Jingyang Chen 외 arxiv

Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied to long-form video understanding scenarios, they exhibit clear limitati…

Reinforcement Learning

Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution

2025-09-28 · Jinpei Guo, Yifei Ji, Shengwei Wang, Zheng Chen 외 arxiv

Diffusion models have recently shown promising results for video super-resolution (VSR). However, directly adapting generative diffusion models to VSR can result in redundancy, since low-quality videos already preserve s…

Video Super-Resolution

An Online Algorithm for Large Scale Image Similarity Learning

2009-12-01 · NeurIPS 2009 12 · Gal Chechik, Uri Shalit, Varun Sharma, Samy Bengio

Learning a measure of similarity between pairs of objects is a fundamental problem in machine learning. It stands in the core of classification methods like kernel machines, and is particularly useful for applications li…

CPU