paper-with-me

홈 › Papers

Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion Models

2024-11-27 · Yiming Wu, Huan Wang, Zhenghao Chen, Dong Xu

The high computational cost and slow inference time are major obstacles to deploying the video diffusion model (VDM) in practical applications. To overcome this, we introduce a new Video Diffusion Model Compression approach using individual content and motion dynamics preserved pruning and consistency loss. First, we empirically observe that deeper VDM layers are crucial for maintaining the quality of \textbf{motion dynamics} e.g., coherence of the entire video, while shallower layers are more focused on \textbf{individual content} e.g., individual frames. Therefore, we prune redundant blocks from the shallower layers while preserving more of the deeper layers, resulting in a lightweight VDM variant called VDMini. Additionally, we propose an \textbf{Individual Content and Motion Dynamics (ICMD)} Consistency Loss to gain comparable generation performance as larger VDM, i.e., the teacher to VDMini i.e., the student. Particularly, we first use the Individual Content Distillation (ICD) Loss to ensure consistency in the features of each generated frame between the teacher and student models. Next, we introduce a Multi-frame Content Adversarial (MCA) Loss to enhance the motion dynamics across the generated video as a whole. This method significantly accelerates inference time while maintaining high-quality video generation. Extensive experiments demonstrate the effectiveness of our VDMini on two important video generation tasks, Text-to-Video (T2V) and Image-to-Video (I2V), where we respectively achieve an average 2.5 $\times$ and 1.4 $\times$ speed up for the I2V method SF-V and the T2V method T2V-Turbo-v2, while maintaining the quality of the generated videos on two benchmarks, i.e., UCF101 and VBench.

📄 PDF Abstract BibTeX arXiv:2411.18375

Code (0)

등록된 구현이 없습니다.

Tasks

Model CompressionVideo Generation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Pruning 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow

2024-12-13 · Zhe Li, Yisheng He, Lei Zhong, Weichao Shen 외

Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style t…

Contrastive LearningMotion Generation

TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

2026-04-09 · Atahan Dokme, Benjamin Reichman, Larry Heck arxiv

Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world queries are often wrapped in frustration, urgency or enthusiasm. Does em…

Measuring Emotional Contagion in Social Media

2015-06-19 · Emilio Ferrara, Zeyao Yang

Social media are used as main discussion channels by millions of individuals every day. The content individuals produce in daily social-media-based micro-communications, and the emotions therein expressed, may impact the…

DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

2023-10-18 · Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen 외

Animating a still image offers an engaging visual experience. Traditional image animation techniques mainly focus on animating natural scenes with stochastic dynamics (e.g. clouds and fluid) or domain-specific motions (e…

Image Animation

Sparsity in Dynamics of Spontaneous Subtle Emotions: Analysis \& Application

2016-01-19 · Anh Cat Le Ngo, John See, Raphael Chung-Wei Phan

Spontaneous subtle emotions are expressed through micro-expressions, which are tiny, sudden and short-lived dynamics of facial muscles; thus poses a great challenge for visual recognition. The abrupt but significant dyna…

Emotion Recognition