paper-with-me

Papers Streaming video understanding

“Streaming video understanding” 태그가 달린 논문 9편 · 필터 해제

InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding

2025-06-18 · Minsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung Chang

Modern multimodal large language models (MLLMs) can reason over hour-long video, yet their key-value (KV) cache grows linearly with time--quickly exceeding the fixed memory of phones, AR glasses, and edge robots. Prior c…

GPUStreaming video understandingTARVideo Understanding

StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant

2025-05-08 · Haibo Wang, Bo Feng, Zhengfeng Lai, Mingze Xu 외

We present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in adapting existing models into online scena…

Language ModelingLanguage ModellingLarge Language ModelStreaming video understanding+1

OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts

2025-03-29 · CVPR 2025 1 · Yuxuan Wang, Yueqian Wang, Bo Chen, Tong Wu 외

The rapid advancement of multi-modal language models (MLLMs) like GPT-4o has propelled the development of Omni language models, designed to process and proactively respond to continuous streams of multi-modal data. Despi…

Streaming video understandingVideo Understanding

ViSpeak: Visual Instruction Feedback in Streaming Videos

2025-03-17 · Shenghao Fu, Qize Yang, Yuan-Ming Li, Yi-Xing Peng 외

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-moda…

Streaming video understandingVideo Understanding

VideoScan: Enabling Efficient Streaming Video Understanding via Frame-level Semantic Carriers

2025-03-12 · Ruanjun Li, Yuedong Tan, Yuanming Shi, Jiawei Shao

This paper introduces VideoScan, an efficient vision-language model (VLM) inference framework designed for real-time video interaction that effectively comprehends and retains streamed video inputs while delivering rapid…

GPUStreaming video understandingVideo Understanding

SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

2025-02-15 · Zhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue 외

Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-cont…

Question AnsweringStreaming video understandingVideo Understanding

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

2025-01-23 · Haomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge 외

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle w…

SchedulingStreaming video understandingVideo Understanding

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

2024-11-06 · Junming Lin, Zheng Fang, Chi Chen, Zihao Wan 외

The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on offline video comprehensi…

Image ComprehensionStreaming video understandingVideo Understanding

System-status-aware Adaptive Network for Online Streaming Video Understanding

2023-03-28 · CVPR 2023 1 · Lin Geng Foo, Jia Gong, Zhipeng Fan, Jun Liu

Recent years have witnessed great progress in deep neural networks for real-time applications. However, most existing works do not explicitly consider the general case where the device's state and the available resources…

Streaming video understandingVideo Understanding
1–9 / 9