paper-with-me

홈 › Papers

Towards Long-Form Video Understanding

2021-06-21 · CVPR 2021 1 · Chao-yuan Wu, Philipp Krähenbühl

Our world offers a never-ending stream of visual stimuli, yet today's vision systems only accurately recognize patterns within a few seconds. These systems understand the present, but fail to contextualize it in past or future events. In this paper, we study long-form video understanding. We introduce a framework for modeling long-form videos and develop evaluation protocols on large-scale datasets. We show that existing state-of-the-art short-term models are limited for long-form tasks. A novel object-centric transformer-based video recognition architecture performs significantly better on 7 diverse tasks. It also outperforms comparable state-of-the-art on the AVA dataset.

📄 PDF Abstract BibTeX arXiv:2106.11310

Code (2)

chaoyuaw/lvu 공식 구현 pytorch
md-mohaiminul/ViS4mer pytorch

Tasks

Action RecognitionFormVideo RecognitionVideo Understanding

Similar Papers 제목 키워드 기반

From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

2024-09-27 · Heqing Zou, Tianze Luo, Guiyang Xie, Victor 외

The integration of Large Language Models (LLMs) with visual encoders has recently shown promising performance in visual understanding tasks, leveraging their inherent capability to comprehend and generate human-like text…

Video UnderstandingVisual Reasoning

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding

2025-01-03 · Heqing Zou, Tianze Luo, Guiyang Xie, Victor Xiao Jie Zhang 외

Multimodal large language models have become a popular topic in deep visual understanding due to many promising real-world applications. However, hour-long video understanding, spanning over one hour and containing tens …

Question AnsweringVideo Understanding

DrVideo: Document Retrieval Based Long Video Understanding

2024-06-18 · CVPR 2025 1 · Ziyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun 외

Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos …

document understandingEgoSchemaMMERetrieval+2

QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models

2026-01-10 · Zixing Lin, Jiale Wang, Gee Wah Ng, Lee Onn Mak 외 arxiv

Large Multimodal Models (LMMs) for video-audio understanding have traditionally been evaluated only on shorter videos of a few minutes long. In this paper, we introduce QMAVIS (Q Team-Multimodal Audio Video Intelligent S…

Speech Recognition

LongVLM: Efficient Long Video Understanding via Large Language Models

2024-04-04 · Yuetian Weng, Mingfei Han, Haoyu He, Xiaojun Chang 외

Empowered by Large Language Models (LLMs), recent advancements in Video-based LLMs (VideoLLMs) have driven progress in various video understanding tasks. These models encode video representations through pooling or query…

Question AnsweringVideo Question AnsweringVideo Understanding