paper-with-me

Papers

Unleashing Hour-Scale Video Training for Long Video-Language Understanding

2025-06-05 · Jingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang, Jiang Liu, Yusheng Su, Xiaodong Yu, Hao Chen, Jiebo Luo, Zicheng Liu, Emad Barsoum

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LLMs underexplored. To close this gap, we present VideoMarathon, a large-scale hour-long video instruction-following dataset. This dataset includes around 9,700 hours of long videos sourced from diverse domains, ranging from 3 to 60 minutes per video. Specifically, it contains 3.3M high-quality QA pairs, spanning six fundamental topics: temporality, spatiality, object, action, scene, and event. Compared to existing video instruction datasets, VideoMarathon significantly extends training video durations up to 1 hour, and supports 22 diverse tasks requiring both short- and long-term video comprehension. Building on VideoMarathon, we propose Hour-LLaVA, a powerful and efficient Video-LMM for hour-scale video-language modeling. It enables hour-long video training and inference at 1-FPS sampling by leveraging a memory augmentation module, which adaptively integrates user question-relevant and spatiotemporal-informative semantics from a cached full video context. In our experiments, Hour-LLaVA achieves the best performance on multiple long video-language benchmarks, demonstrating the high quality of the VideoMarathon dataset and the superiority of the Hour-LLaVA model.

📄 PDF Abstract BibTeX arXiv:2506.05332

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding

2025-01-03 · Heqing Zou, Tianze Luo, Guiyang Xie, Victor Xiao Jie Zhang 외

Multimodal large language models have become a popular topic in deep visual understanding due to many promising real-world applications. However, hour-long video understanding, spanning over one hour and containing tens …

Question AnsweringVideo Understanding

Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition

2026-06-10 · Sukmin Seo, Geewook Kim arxiv

Temporal grounding--returning the interval $[t_s, t_e]$ for a natural-language query over a video--is the language interface to long-form video, yet has been studied on short videos; the dynamics of hour-scale natural-la…

SVHighlights: Towards Extremely Long Sport Video Highlight Detection

2026-06-05 · Donggyu Lee, Youngbin Ki, Jeonghun Kang, Taehwan Kim arxiv

While highlight detection for long-form videos is of great practical importance, most existing methods remain limited to short-form content, largely due to the absence of a suitable benchmark. To bridge this gap, we intr…

Highlight Detection

Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors

2026-06-26 · Youzhi Liu, Li Gao, Yifei Qian, Liu Liu 외 arxiv

Quadruped robots have achieved remarkable locomotion, yet their behavioral repertoire remains confined to a few gaits--far from the expressive, companion-like presence long envisioned for them. Attempts to import the hum…

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

2025-05-29 · David Ma, Huaqing Yuan, Xingjian Wang, Qianbo Zang 외

Although long-video understanding demands that models capture hierarchical temporal information -- from clip (seconds) and shot (tens of seconds) to event (minutes) and story (hours) -- existing benchmarks either neglect…

AvgVideo Understanding