paper-with-me

홈 › Papers

Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video

2025-10-16 · Yulin Zhang, Cheng Shi, Yang Wang, Sibei Yang arxiv

Envision an AI capable of functioning in human-like settings, moving beyond mere observation to actively understand, anticipate, and proactively respond to unfolding events. Towards this vision, we focus on the innovative task where, given ego-streaming video input, an assistant proactively answers diverse, evolving questions at the opportune moment, while maintaining synchronized perception and reasoning. This task embodies three key properties: (1) Proactive Coherence, (2) Just-in-Time Responsiveness, and (3) Synchronized Efficiency. To evaluate and address these properties, we first introduce ESTP-Bench (Ego Streaming Proactive Benchmark) alongside the ESTP-F1 metric-a novel framework designed for their rigorous assessment. Secondly, we propose a comprehensive technical pipeline to enable models to tackle this challenging task. This pipeline comprises: (1) a data engine, (2) a multi-stage training strategy, and (3) a proactive dynamic compression technique. Our proposed model effectively addresses these critical properties while outperforming multiple baselines across diverse online and offline benchmarks. Project Page:https://zhangyl4.github.io/publications/eyes-wide-open/

📄 PDF Abstract BibTeX arXiv:2510.14560

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding

2026-03-19 · Yikai Zheng, Xin Ding, Yifan Yang, Shiqi Jiang 외 arxiv

Recent advances in Streaming Video Understanding has enabled a new interaction paradigm where models respond proactively to user queries. Current proactive VideoLLMs rely on per-frame triggering decision making, which su…

Decision Making

Harnessing Streaming Video in the Wild

2026-06-07 · Dingyu Yao, Shuhuan Gu, Qingyi Si, Junhao Zhou 외 arxiv

Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentary, and embodied robots. An ideal streaming system should support pro…

StreamingClaw Technical Report

2026-03-23 · Jiawei Chen, Zhe Chen, Chaoqun Du, Maokui He 외 arxiv

Emerging applications such as embodied intelligence, AI hardware, autonomous driving, and intelligent cockpits rely on a real-time perception-decision-action closed loop, posing stringent challenges for streaming video u…

Autonomous Driving

MOSS-VL Technical Report

2026-08-15 · Pengyu Wang, Chenkun Tan, Shaojun Zhou, Qirui Zhou 외 hf

We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to…

STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding

2026-03-29 · Junho Kim, Hosu Lee, James M. Rehg, Minsu Kim 외 arxiv

Recent progress in video large language models (Video-LLMs) has enabled strong offline reasoning over long and complex videos. However, real-world deployments increasingly require streaming perception and proactive inter…