paper-with-me

홈 › Papers

AdaCM^2: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction

2025-01-01 · CVPR 2025 1 · Yuanbin Man, Ying Huang, Chengming Zhang, Bingzhe Li, Wei Niu, Miao Yin

The advancements in large language models (LLMs) have propelled the improvement of video understanding tasks by incorporating LLMs with visual models. However, most existing LLM-based models (e.g., VideoLLaMA, VideoChat) are constrained to processing short-duration videos. Recent attempts to understand long-term videos by extracting and compressing visual features into a fixed memory size. Nevertheless, those methods leverage only visual modality to merge video tokens and overlook the correlation between visual and textual queries, leading to difficulties in effectively handling complex question-answering tasks. To address the challenges of long videos and complex prompts, we propose AdaCM^2, which, for the first time, introduces an adaptive cross-modality memory reduction approach to video-text alignment in an auto-regressive manner on video streams. Our extensive experiments on various video understanding tasks, such as video captioning, video question answering, and video classification, demonstrate that AdaCM^2 achieves state-of-the-art performance across multiple datasets while significantly reducing memory usage. Notably, it achieves a 4.5% improvement across multiple tasks in the LVU dataset with a GPU memory consumption reduction of up to 65%.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

GPUQuestion AnsweringVideo CaptioningVideo ClassificationVideo Question AnsweringVideo Understanding

Similar Papers 제목 키워드 기반

AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction

2024-11-19 · Yuanbin Man, Ying Huang, Chengming Zhang, Bingzhe Li 외

The advancements in large language models (LLMs) have propelled the improvement of video understanding tasks by incorporating LLMs with visual models. However, most existing LLM-based models (e.g., VideoLLaMA, VideoChat)…

GPUQuestion AnsweringVideo CaptioningVideo Classification+2

AdaCM: Adaptive ColorMLP for Real-Time Universal Photo-realistic Style Transfer

2022-12-03 · Tianwei Lin, Honglin Lin, Fu Li, Dongliang He 외

Photo-realistic style transfer aims at migrating the artistic style from an exemplar style image to a content image, producing a result image without spatial distortions or unrealistic artifacts. Impressive results have …

4kGPUStyle Transfer

X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding

2025-01-12 · Wenqi Zhou, Kai Cao, Hao Zheng, Xinyi Zheng 외

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activit…

Video Understanding

Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

2024-06-12 · Haoji Zhang, Yiqin Wang, Yansong Tang, Yong liu 외

Benefiting from the advancements in large language models and cross-modal alignment, existing multi-modal video understanding methods have achieved prominent performance in offline scenario. However, online video streams…

cross-modal alignmentLanguage ModellingQuestion AnsweringVideo Question Answering+2

Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

2025-03-24 · Xiangrui Liu, Yan Shu, Zheng Liu, Ao Li 외

Despite advanced token compression techniques, existing multimodal large language models (MLLMs) still struggle with hour-long video understanding. In this work, we propose Video-XL-Pro, an efficient method for extremely…

8kGPUSelf-Supervised LearningVideo Understanding