paper-with-me

홈 › Papers

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding

2026-04-03 · Yufei Yin, Yuchen Xing, Qianke Meng, Minghao Chen, Yan Yang, Zhou Yu arxiv

Understanding long videos requires extracting query-relevant information from long sequences under tight compute budgets. Existing text-then-LLM pipelines lose fine-grained visual cues, while video-based multimodal large language models (MLLMs) can keep visual details but are too frame-hungry and computationally expensive. In this work, we aim to harness MLLMs for efficient video understanding. We propose ProVCA, a progressive video condensation agent that iteratively locates key video frames at multiple granularities. ProVCA first adopts a segment localization module to identify the video segment relevant to the query, then a snippet selection module to select important snippets based on similarity, and finally a keyframe refinement module to pinpoint specific keyframes in those snippets. By progressively narrowing the scope from coarse segments to fine frames, ProVCA identifies a small set of keyframes for MLLM-based reasoning. ProVCA achieves state-of-the-art zero-shot accuracies of 69.3\% on EgoSchema, 80.5\% on NExT-QA, and 77.7\% on IntentQA, while using fewer frames than previous training-free methods.

📄 PDF Abstract BibTeX arXiv:2604.02891

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents

2025-03-13 · BoYu Chen, Zhengrong Yue, Siran Chen, Zikang Wang 외

Existing Multimodal Large Language Models (MLLMs) encounter significant challenges in modeling the temporal context within long videos. Currently, mainstream Agent-based methods use external tools (e.g., search engine, m…

Computational EfficiencyOptical Character Recognition (OCR)RetrievalVideo Understanding

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

2026-05-01 · Kerui Chen, Jinglu Wang, Jianrong Zhang, Ming Li 외 arxiv

Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existing agentic methods mitigate this via rule…

VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning

2025-12-26 · Yang Ding, Yizhen Zhang, Xin Lai, Ruihang Chu 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language tasks yet remain limited in long video understanding due to the limited context window. Consequently, prevailing approaches te…

Reinforcement Learning

GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory

2025-11-15 · Jeong Hun Yeo, Sangyun Chung, Sungjune Park, Dae Hoe Kim 외 arxiv

Long-video understanding remains a significant challenge for Multimodal Large Language Models (MLLMs) due to inherent token limitations and the complexity of capturing long-term temporal dependencies. Existing methods of…

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

2026-07-23 · Lihuang Fang, Yuchen Zou, kebin Jin, Jinghui Qin arxiv

Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion understanding with advanced video understa…

Multimodal Emotion RecognitionReinforcement Learning