paper-with-me

홈 › Papers

Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents

2025-09-29 · Jiahua Li, Zhanhe Zhang, Chenghao Xu, Zhe Xu, Kun Wei, Xu Yang, Cheng Deng arxiv

Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model (LLM)-based approaches have advanced long video understanding, they remain bottlenecked by task-agnostic, fixed-granularity perception pipelines and suffer from vision-language hallucinations. Inspired by human adaptive perception and active verification, we propose CogniGPT, a framework leveraging an interactive loop between a Multi-Granular Perception Agent (MPA) and an Active Verification Agent (AVA). Specifically, instead of predetermined heuristics, MPA adaptively determines the optimal perception granularity and strategy based on the evolving context, while AVA actively mines multi-perspective visual evidence to cross-verify key observations and eliminate hallucinations. This interaction allows CogniGPT to efficiently identify a minimal set of reliable task-related clues. Extensive experiments on EgoSchema, Video-MME, NExT-QA, and MovieChat demonstrate its superiority in accuracy and efficiency. Notably, on EgoSchema, it surpasses existing training-free methods using only 11.2 frames and achieves performance comparable to Gemini 1.5-Pro.

📄 PDF Abstract BibTeX arXiv:2509.24943

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models

2025-11-24 · Fufangchen Zhao, Liao Zhang, Daiqi Shi, Yuanjun Gao 외 arxiv

We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to reason about brief actions in short cli…

Reinforcement LearningAction Understanding

VideoDeepResearch: Long Video Understanding With Agentic Tool Using

2025-06-12 · Huaying Yuan, Zheng Liu, Junjie Zhou, Ji-Rong Wen 외

Long video understanding (LVU) presents a significant challenge for current multi-modal large language models (MLLMs) due to the task's inherent complexity and context window constraint. It is widely assumed that address…

MMEVideo MMEVideo Understanding

MM-AU:Towards Multimodal Understanding of Advertisement Videos

2023-08-27 · Digbalay Bose, Rajat Hebbar, Tiantian Feng, Krishna Somandepalli 외

Advertisement videos (ads) play an integral part in the domain of Internet e-commerce as they amplify the reach of particular products to a broad audience or can serve as a medium to raise awareness about specific issues…

iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering

2020-11-16 · Aman Chadha, Gurneet Arora, Navpreet Kaloty

Most prior art in visual understanding relies solely on analyzing the "what" (e.g., event recognition) and "where" (e.g., event localization), which in some cases, fails to describe correct contextual relationships betwe…

Common Sense ReasoningDense Video CaptioningMachine TranslationQuestion Answering+2

AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

2025-03-16 · Xiao Wang, Qingyi Si, Jianlong Wu, Shiyu Zhu 외

Multimodal Large Language Models (MLLMs) have revolutionized video understanding, yet are still limited by context length when processing long videos. Recent methods compress videos by leveraging visual redundancy unifor…

Video Understanding