paper-with-me

홈 › Papers

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

2026-06-08 · Apratim Bhattacharyya, Shweta Mahajan, Sanjay Haresh, Rajeev Yasarla, Reza Pourreza, Litian Liu, Risheek Garrepalli, Roland Memisevic arxiv

Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large language models (LLMs) as task guidance assistants. A crucial capability for the real-world success of a prospective task guidance assistant is it's ability to intervene proactively as soon as a mistake is apparent in order to guide the user. To evaluate this crucial capability, we introduce Ego-MC-Bench (Mistake Corrections), a benchmark for evaluating reactive, step-by-step task guidance in realistic cooking scenarios. Extensive experiments show that Ego-MC-Bench is highly challenging for state-of-the-art video LLMs. We argue that a key reason is the limited availability of training data for fine-tuning models on this task. Although there exists a wide range of cooking video datasets, existing datasets lack examples of mistakes along with appropriately timed interventions. To help address this data limitation, we also introduce Ego-CoMist, a counterfactual synthetic dataset created by transforming non -interactive cooking videos into supervised training examples showing proactive interventions. We show that fine-tuning on Ego-CoMist yields performance gains especially for smaller and more efficient video LLMs that are well suited for delivering assistance on edge devices.

📄 PDF Abstract BibTeX arXiv:2606.09547

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

2026-06-29 · Shuchang Ye, Jinqiang Yu, Zhujun Xiao, Yajing Kong 외 arxiv

Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation to platform-specific data distributions, …

VideoLLM-online: Online Video Large Language Model for Streaming Video

2024-06-17 · CVPR 2024 1 · Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin 외

Recent Large Language Models have been enhanced with vision capabilities, enabling them to comprehend images, videos, and interleaved vision-language content. However, the learning methods of these large multimodal model…

GPULanguage ModelingLanguage ModellingLarge Language Model+1

LiveStar: Live Streaming Assistant for Real-World Online Video Understanding

2025-11-07 · Zhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang 외 arxiv

Despite significant progress in Video Large Language Models (Video-LLMs) for offline video understanding, existing online Video-LLMs typically struggle to simultaneously process continuous frame-by-frame inputs and deter…

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams

2026-06-16 · Zhenyu Yang, Kairui Zhang, Bing Wang, Shengsheng Qian 외 arxiv

Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preser…

BehanceMT: A Machine Translation Corpus for Livestreaming Video Transcripts

2022-10-01 · TU (COLING) 2022 10 · Minh Van Nguyen, Franck Dernoncourt, Thien Nguyen

Machine translation (MT) is an important task in natural language processing, which aims to translate a sentence in a source language to another sentence with the same/similar semantics in a target language. Despite the …

ArticlesMachine TranslationSentenceTranslation