paper-with-me

홈 › Papers

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

2026-07-04 · Xuyao Huang, Zelai Deng, Xu Wang, Xizhong Xiao, Zhijie Deng hf

The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterative denoising processes lead to substantial computational overhead, making efficient generation on mobile devices challenging. We propose CineMobile to bridge the gap. In particular, CineMobile adopts a three-fold optimization strategy: (1) leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects; (2) optimizing the compressed model into a 4-step generator via a combination of diffusion distillation and reinforcement learning; (3) employing a hybrid post-training quantization strategy to compress the model footprint to under 1 GB. Experimental results show that compared to the teacher model with the Wan 2.1 architecture, CineMobile achieves a 40x speedup in generation while maintaining comparable visual quality. Specifically, CineMobile generates 49-frame 480p videos with a per-step denoising latency of 0.6s on an NVIDIA H200 GPU and 20s on the MediaTek Dimensity 8400 Ultimate 5G platform, with a peak memory usage of 1.8 GB, demonstrating its practical applicability for mobile-based image-to-video creation.

📄 PDF Abstract BibTeX arXiv:2607.03803

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningVideo Generation

Similar Papers 제목 키워드 기반

Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models

2025-04-06 · Sridhar S, Nithin A, Shakeel Rifath, Vasantha Raj

Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies inc…

Audio GenerationGPUImage GenerationVideo Synchronization

Can video generation replace cinematographers? Research on the cinematic language of generated video

2024-12-16 · Xiaozhe Li, Kai Wu, Siyi Yang, YiZhan Qu 외

Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing research primarily focuses on object mo…

Video Generation

ShotPlan: Cinematic Video Generation with Learnable Planning Token

2026-07-20 · Su Guo, Guangce Liu, Haosen Yang, Jiepeng Wang 외 hf

Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot…

Video Generation

SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

2024-12-13 · CVPR 2025 1 · Yushu Wu, Zhixing Zhang, Yanyu Li, Yanwu Xu 외

We have witnessed the unprecedented success of diffusion-based video generation over the past year. Recently proposed models from the community have wielded the power to generate cinematic and high-resolution videos with…

DenoisingImage GenerationVideo Generation

CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models

2025-08-15 · Xiaoxue Wu, Bingjie Gao, Yu Qiao, Yaohui Wang 외 arxiv

Despite significant advances in video synthesis, research into multi-shot video generation remains in its infancy. Even with scaled-up models and massive datasets, the shot transition capabilities remain rudimentary and …

Video Generation