paper-with-me

Papers

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding

2025-01-03 · Heqing Zou, Tianze Luo, Guiyang Xie, Victor Xiao Jie Zhang, Fengmao Lv, Guangcong Wang, Junyang Chen, Zhuochen Wang, Hansheng Zhang, Huaijian Zhang

Multimodal large language models have become a popular topic in deep visual understanding due to many promising real-world applications. However, hour-long video understanding, spanning over one hour and containing tens of thousands of visual frames, remains under-explored because of 1) challenging long-term video analyses, 2) inefficient large-model approaches, and 3) lack of large-scale benchmark datasets. Among them, in this paper, we focus on building a large-scale hour-long long video benchmark, HLV-1K, designed to evaluate long video understanding models. HLV-1K comprises 1009 hour-long videos with 14,847 high-quality question answering (QA) and multi-choice question asnwering (MCQA) pairs with time-aware query and diverse annotations, covering frame-level, within-event-level, cross-event-level, and long-term reasoning tasks. We evaluate our benchmark using existing state-of-the-art methods and demonstrate its value for testing deep long video understanding capabilities at different levels and for various tasks. This includes promoting future long video understanding tasks at a granular level, such as deep understanding of long live videos, meeting recordings, and movies.

📄 PDF Abstract BibTeX arXiv:2501.01645

Code (1)

vincent-zhq/hlv-1k 공식 구현

Tasks

Question AnsweringVideo Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Unleashing Hour-Scale Video Training for Long Video-Language Understanding

2025-06-05 · Jingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang 외

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LL…

Instruction FollowingLanguage ModelingLanguage Modelling

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition

2026-06-10 · Sukmin Seo, Geewook Kim arxiv

Temporal grounding--returning the interval $[t_s, t_e]$ for a natural-language query over a video--is the language interface to long-form video, yet has been studied on short videos; the dynamics of hour-scale natural-la…

SVHighlights: Towards Extremely Long Sport Video Highlight Detection

2026-06-05 · Donggyu Lee, Youngbin Ki, Jeonghun Kang, Taehwan Kim arxiv

While highlight detection for long-form videos is of great practical importance, most existing methods remain limited to short-form content, largely due to the absence of a suitable benchmark. To bridge this gap, we intr…

Highlight Detection

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

2025-05-29 · David Ma, Huaqing Yuan, Xingjian Wang, Qianbo Zang 외

Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours) -- existing benchmarks either neglect…

AvgVideo Understanding

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

2025-03-31 · CVPR 2025 1 · Lucas Ventura, Antoine Yang, Cordelia Schmid, Gül Varol

We address the task of video chaptering, i.e., partitioning a long video timeline into semantic units and generating corresponding chapter titles. While relatively underexplored, automatic chaptering has the potential to…

Large Language ModelVideo Chaptering