paper-with-me

홈 › Papers

Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding

2026-07-14 · Yifan Lu, Ziqi Zhang, Chunfeng Yuan, Jun Gao, Bing Li, Weiming Hu arxiv

Large Vision-Language Models (LVLMs) face significant challenges in long video understanding due to the excessive computational cost and information loss associated with uniform sampling. Existing keyframe selection methods often treat video frames as atomic entities and allocate visual budgets equally, thereby overlooking high-level semantic structures and introducing substantial redundancy. To address these limitations, we propose GMM-EVA (Gaussian Mixture Modeling for Event-Aware Visual Allocation), which leverages Gaussian Mixture Models to model event-level structure from discrete frame-wise observations. A differentiated allocation strategy is then applied to preserve one primary high-resolution keyframe per event for high-fidelity detail, while utilizing lower-resolution secondary keyframes to maintain temporal context and optimize token budgets. GMM-EVA is a training-free, plug-and-play framework that generalizes robustly across various relevance measures and downstream LVLMs. Extensive experiments on multiple long video benchmarks demonstrate that our method significantly outperforms uniform sampling. Notably, GMM-EVA achieves comparable performance to baseline selection methods while utilizing only approximately half of the visual token budget, highlighting its superior efficiency and effectiveness.

📄 PDF Abstract BibTeX arXiv:2607.12557

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization

2025-07-14 · Huilai Li, Yonghao Dang, Ying Xing, Yiming Wang 외 arxiv

Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. Most studies only employ event-related semantic constraints on the final outputs, lac…

audio-visual event localization

Visual Sensor Network Stimulation Model Identification via Gaussian Mixture Model and Deep Embedded Features

2022-01-18 · Luca Varotto, Marco Fabris, Giulia Michieletto, Angelo Cenedese

Visual sensor networks (VSNs) constitute a fundamental class of distributed sensing systems, with unique complexity and appealing performance features, which correspondingly bring in quite active lines of research. An im…

model

Ordering-sensitive and Semantic-aware Topic Modeling

2015-02-12 · Min Yang, Tianyi Cui, Wenting Tu

Topic modeling of textual corpora is an important and challenging problem. In most previous work, the "bag-of-words" assumption is usually made which ignores the ordering of words. This assumption simplifies the computat…

Retrieval

How do Mixture Density RNNs Predict the Future?

2019-01-23 · Kai Olav Ellefsen, Charles Patrick Martin, Jim Torresen

Gaining a better understanding of how and what machine learning systems learn is important to increase confidence in their decisions and catalyze further research. In this paper, we analyze the predictions made by a spec…

Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling

2025-11-11 · Jiale Liu, Haoming Zhou, Yishu Liu, Bingzhi Chen 외 arxiv

Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment,…

Vision-Language NavigationVisual Question AnsweringImage Captioning