paper-with-me

Papers

Infinite Video Understanding

2025-07-11 · Dell Zhang, Xiangyu Chen, Jixiang Luo, Mengxi Jia, Changzhi Sun, Ruilong Ren, Jingren Liu, Hao Sun, Xuelong Li arxiv

The rapid advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have ushered in remarkable progress in video understanding. However, a fundamental challenge persists: effectively processing and comprehending video content that extends beyond minutes or hours. While recent efforts like Video-XL-2 have demonstrated novel architectural solutions for extreme efficiency, and advancements in positional encoding such as HoPE and VideoRoPE++ aim to improve spatio-temporal understanding over extensive contexts, current state-of-the-art models still encounter significant computational and memory constraints when faced with the sheer volume of visual tokens from lengthy sequences. Furthermore, maintaining temporal coherence, tracking complex events, and preserving fine-grained details over extended periods remain formidable hurdles, despite progress in agentic reasoning systems like Deep Video Discovery. This position paper posits that a logical, albeit ambitious, next frontier for multimedia research is Infinite Video Understanding -- the capability for models to continuously process, understand, and reason about video data of arbitrary, potentially never-ending duration. We argue that framing Infinite Video Understanding as a blue-sky research objective provides a vital north star for the multimedia, and the wider AI, research communities, driving innovation in areas such as streaming architectures, persistent memory mechanisms, hierarchical and adaptive representations, event-centric reasoning, and novel evaluation paradigms. Drawing inspiration from recent work on long/ultra-long video understanding and several closely related fields, we outline the core challenges and key research directions towards achieving this transformative capability.

📄 PDF Abstract BibTeX arXiv:2507.09068

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

StreamingVLM: Real-Time Understanding for Infinite Video Streams

2025-10-10 · Ruyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He 외 arxiv

Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite video streams without escalating latency and memory usage. Processing …

Towards Smooth Video Composition

2022-12-14 · Qihang Zhang, Ceyuan Yang, Yujun Shen, Yinghao Xu 외

Video generation requires synthesizing consistent and persistent frames with dynamic content over time. This work investigates modeling the temporal relations for composing video with arbitrary length, from a few frames …

Image Generationsingle-image-generationVideo GenerationVideo Understanding

Infinite Gaze Generation for Videos with Autoregressive Diffusion

2026-03-26 · Jenna Kang, Colin Groth, Tong Wu, Finley Torrens 외 arxiv

Predicting human gaze in video is fundamental to advancing scene understanding and multimodal interaction. While traditional saliency maps provide spatial probability distributions and scanpaths offer ordered fixations, …

Scene Understanding

Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism

2026-03-31 · Tao Chen, Kun Zhang, Qiong Wu, Xiao Chen 외 arxiv

Long video understanding is a key challenge that plagues the advancement of \emph{Multimodal Large language Models} (MLLMs). In this paper, we study this problem from the perspective of visual memory mechanism, and propo…

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation

2025-08-11 · Shuyuan Tu, Yueming Pan, Yinming Huang, Xintong Han 외 arxiv

Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end vid…

Video Generation