paper-with-me

홈 › Papers

QVGen: Pushing the Limit of Quantized Video Generative Models

2025-05-16 · Yushi Huang, Ruihao Gong, Jing Liu, Yifu Ding, Chengtao Lv, Haotong Qin, Jun Zhang

Video diffusion models (DMs) have enabled high-quality video synthesis. Yet, their substantial computational and memory demands pose serious challenges to real-world deployment, even on high-end GPUs. As a commonly adopted solution, quantization has proven notable success in reducing cost for image DMs, while its direct application to video DMs remains ineffective. In this paper, we present QVGen, a novel quantization-aware training (QAT) framework tailored for high-performance and inference-efficient video DMs under extremely low-bit quantization (e.g., 4-bit or below). We begin with a theoretical analysis demonstrating that reducing the gradient norm is essential to facilitate convergence for QAT. To this end, we introduce auxiliary modules ($\Phi$) to mitigate large quantization errors, leading to significantly enhanced convergence. To eliminate the inference overhead of $\Phi$, we propose a rank-decay strategy that progressively eliminates $\Phi$. Specifically, we repeatedly employ singular value decomposition (SVD) and a proposed rank-based regularization $\mathbf{\gamma}$ to identify and decay low-contributing components. This strategy retains performance while zeroing out inference overhead. Extensive experiments across $4$ state-of-the-art (SOTA) video DMs, with parameter sizes ranging from $1.3$B $\sim14$B, show that QVGen is the first to reach full-precision comparable quality under 4-bit settings. Moreover, it significantly outperforms existing methods. For instance, our 3-bit CogVideoX-2B achieves improvements of $+25.28$ in Dynamic Degree and $+8.43$ in Scene Consistency on VBench.

📄 PDF Abstract BibTeX arXiv:2505.11497

Code (0)

등록된 구현이 없습니다.

Tasks

Quantization

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Video Prediction Models as General Visual Encoders

2024-05-25 · James Maier, Nishanth Mohankumar

This study explores the potential of open-source video conditional generation models as encoders for downstream tasks, focusing on instance segmentation using the BAIR Robot Pushing Dataset. The researchers propose using…

Instance SegmentationPredictionSegmentationSemantic Segmentation+1

RefTok: Reference-Based Tokenization for Video Generation

2025-07-03 · Xiang Fan, Xiaohang Sun, Kushan Thakkar, Zhu Liu 외 arxiv

Effectively handling temporal redundancy remains a key challenge in learning video models. Prevailing approaches often treat each set of frames independently, failing to effectively capture the temporal dependencies and …

Video Generation

EdiBERT, a generative model for image editing

2021-11-30 · Thibaut Issenhuth, Ugo Tanielian, Jérémie Mary, David Picard

Advances in computer vision are pushing the limits of im-age manipulation, with generative models sampling detailed images on various tasks. However, a specialized model is often developed and trained for each specific t…

DenoisingImage DenoisingImage Manipulationmodel

QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks

2026-05-21 · Haotong Qin, Xudong Ma, Xianglong Liu, Jie Luo 외 arxiv

Low-bit quantization is widely used to compress super-resolution (SR) models and reduce storage and computation costs for deployment on resource-limited devices. However, when SR models are pushed to ultra-low precision …

Image Super-Resolution

MambaVideo for Discrete Video Tokenization with Channel-Split Quantization

2025-07-06 · Dawit Mureja Argaw, Xian Liu, Joon Son Chung, Ming-Yu Liu 외 arxiv

Discrete video tokenization is essential for efficient autoregressive generative modeling due to the high dimensionality of video data. This work introduces a state-of-the-art discrete video tokenizer with two key contri…

Video Generation