paper-with-me

홈 › Papers

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

2025-09-29 · Junsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu, Junyu Chen, Shuai Yang, Xianbang Wang, Yicheng Pan, Daquan Zhou, Huan Ling, Haozhe Liu, Hongwei Yi, Hao Zhang, Muyang Li, Yukang Chen, Han Cai, Sanja Fidler, Ping Luo, Song Han, Enze Xie arxiv

We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720x1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong text-video alignment at a remarkably fast speed, deployable on RTX 5090 GPU. Two core designs ensure our efficient, effective and long video generation: (1) Linear DiT: We leverage linear attention as the core operation, which is more efficient than vanilla attention given the large number of tokens processed in video generation. (2) Constant-Memory KV cache for Block Linear Attention: we design block-wise autoregressive approach for long video generation by employing a constant-memory state, derived from the cumulative properties of linear attention. This KV cache provides the Linear DiT with global context at a fixed memory cost, eliminating the need for a traditional KV cache and enabling efficient, minute-long video generation. In addition, we explore effective data filters and model training strategies, narrowing the training cost to 12 days on 64 H100 GPUs, which is only 1% of the cost of MovieGen. Given its low cost, SANA-Video achieves competitive performance compared to modern state-of-the-art small diffusion models (e.g., Wan 2.1-1.3B and SkyReel-V2-1.3B) while being 16x faster in measured latency. Moreover, SANA-Video can be deployed on RTX 5090 GPUs with NVFP4 precision, accelerating the inference speed of generating a 5-second 720p video from 71s to 29s (2.4x speedup). In summary, SANA-Video enables low-cost, high-quality video generation.

📄 PDF Abstract BibTeX arXiv:2509.24695

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationVideo Alignment

Similar Papers 제목 키워드 기반

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

2026-07-23 · Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue 외 arxiv

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches…

Video Generation

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

2026-05-14 · Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye 외 arxiv

We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p, minute-scale videos with precise camera control. SANA-WM achieves vi…

SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

2026-05-28 · Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu 외 arxiv

Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formidable challenge due to the stringent requirements for temporal consist…

ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers

2026-01-07 · Mohsen Ghafoorian, Amirhossein Habibian arxiv

Recent advances in video diffusion models have shifted towards transformer-based architectures, achieving state-of-the-art video generation but at the cost of quadratic attention complexity, which severely limits scalabi…

Video Generation

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

2025-01-30 · Enze Xie, Junsong Chen, Yuyang Zhao, Jincheng Yu 외

This paper presents SANA-1.5, a linear Diffusion Transformer for efficient scaling in text-to-image generation. Building upon SANA-1.0, we introduce three key innovations: (1) Efficient Training Scaling: A depth-growth p…

Image GenerationModel CompressionText to Image GenerationText-to-Image Generation