paper-with-me

Papers

MTC-VAE: Multi-Level Temporal Compression with Content Awareness

2026-02-01 · Yubo Dong, Linchao Zhu arxiv

Latent Video Diffusion Models (LVDMs) rely on Variational Autoencoders (VAEs) to compress videos into compact latent representations. For continuous Variational Autoencoders (VAEs), achieving higher compression rates is desirable; yet, the efficiency notably declines when extra sampling layers are added without expanding the dimensions of hidden channels. In this paper, we present a technique to convert fixed compression rate VAEs into models that support multi-level temporal compression, providing a straightforward and minimal fine-tuning approach to counteract performance decline at elevated compression rates.Moreover, we examine how varying compression levels impact model performance over video segments with diverse characteristics, offering empirical evidence on the effectiveness of our proposed approach. We also investigate the integration of our multi-level temporal compression VAE with diffusion-based generative models, DiT, highlighting successful concurrent training and compatibility within these frameworks. This investigation illustrates the potential uses of multi-level temporal compression.

📄 PDF Abstract BibTeX arXiv:2602.01340

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation

2026-07-03 · Syed Ariff Syed Hesham, Yun Liu, Guolei Sun, Jing Yang 외 arxiv

Video reasoning segmentation demands pixel-accurate object tracking across hundreds of frames under complex natural language queries, producing dense spatiotemporal tokens whose quadratic self-attention cost makes long-v…

Natural Language QueriesObject Tracking

Will You Be Aware? Eye Tracking-Based Modeling of Situational Awareness in Augmented Reality

2025-08-07 · Zhehan Qu, Tianyi Hu, Christian Fronk, Maria Gorlatova arxiv

Augmented Reality (AR) systems, while enhancing task performance through real-time guidance, pose risks of inducing cognitive tunneling-a hyperfocus on virtual content that compromises situational awareness (SA) in safet…

Graph Neural Network

StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models

2024-08-31 · Yuxiang Guo, Faizan Siddiqui, Yang Zhao, Rama Chellappa 외

Predicting and reasoning how a video would make a human feel is crucial for developing socially intelligent systems. Although Multimodal Large Language Models (MLLMs) have shown impressive video understanding capabilitie…

Video Understanding

EA-VTR: Event-Aware Video-Text Retrieval

2024-07-10 · Zongyang Ma, Ziqi Zhang, Yuxin Chen, Zhongang Qi 외

Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and th…

Action RecognitionContrastive Learningcross-modal alignmentMoment Retrieval+6

InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models

2026-06-01 · Xinxin Liu, Shiwei Gan, Xiao Liu, Yafeng Yin 외 arxiv

Video Large Language Models (Video-LLMs) achieve strong performance in video understanding, but their excessive visual tokens bring substantial computational overhead. Existing training-free compression methods improve i…