paper-with-me

Papers

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

2025-09-15 · Hyolim Kang, Yunsu Park, Youngbeom Yoo, Yeeun Choi, Seon Joo Kim arxiv

We introduce Hierarchical Streaming Video Understanding, a task that combines online temporal action localization with free-form description generation. Given the scarcity of datasets with hierarchical and fine-grained temporal annotations, we demonstrate that LLMs can effectively group atomic actions into higher-level events, enriching existing datasets. We then propose OpenHOUSE (Open-ended Hierarchical Online Understanding System for Events), which extends streaming action perception beyond action classification. OpenHOUSE features a specialized streaming module that accurately detects boundaries between closely adjacent actions, nearly doubling the performance of direct extensions of existing methods. We envision the future of streaming action perception in the integration of powerful generative models, with OpenHOUSE representing a key step in that direction.

📄 PDF Abstract BibTeX arXiv:2509.12145

Code (0)

등록된 구현이 없습니다.

Tasks

Temporal Action LocalizationAction Classification

Similar Papers 제목 키워드 기반

H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding

2025-03-31 · Qi Wu, Quanlong Zheng, Yanhao Zhang, Junlin Xie 외

With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating video understanding exhibit significant…

Video Understanding

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

2025-01-23 · Haomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge 외

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle w…

SchedulingStreaming video understandingVideo Understanding

StreamingClaw Technical Report

2026-03-23 · Jiawei Chen, Zhe Chen, Chaoqun Du, Maokui He 외 arxiv

Emerging applications such as embodied intelligence, AI hardware, autonomous driving, and intelligent cockpits rely on a real-time perception-decision-action closed loop, posing stringent challenges for streaming video u…

Autonomous Driving

AURA: Always-On Understanding and Real-Time Assistance via Video Streams

2026-04-05 · Xudong Lu, Yang Bo, Jinpeng Chen, Shuhan Li 외 arxiv

Video Large Language Models (VideoLLMs) have achieved strong performance on many video understanding tasks, but most existing systems remain offline and are not well-suited for live video streams that require continuous …

Question Answering

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

2026-01-21 · Haowei Zhang, Shudong Yang, Jinlan Fu, See-Kiong Ng 외 arxiv

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated significant improvement in offline video understanding. However, extending these capabilities to streaming video inputs, remains challengi…