paper-with-me

홈 › Papers

SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

2026-05-14 · Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han, Enze Xie arxiv

We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p, minute-scale videos with precise camera control. SANA-WM achieves visual quality comparable to large-scale industrial baselines such as LingBot-World and HY-WorldPlay, while significantly improving efficiency. Four core designs drive our architecture: (1) Hybrid Linear Attention combines frame-wise Gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling. (2) Dual-Branch Camera Control ensures precise 6-DoF trajectory adherence. (3) Two-Stage Generation Pipeline applies a long-video refiner to stage-1 outputs, improving quality and consistency across sequences. (4) Robust Annotation Pipeline extracts accurate metric-scale 6-DoF camera poses from public videos to yield high-quality, spatiotemporally consistent action labels. Driven by these designs, SANA-WMdemonstrates remarkable efficiency across data, training compute, and inference hardware: it uses only $\sim$213K public video clips with metric-scale pose supervision, completes training in 15 days on 64 H100s, and generates each 60s clip on a single GPU; its distilled variant can be deployed on a single RTX 5090 with NVFP4 quantization to denoise a 60s 720p clip in 34s. On our one-minute world-model benchmark, SANA-WM demonstrates stronger action-following accuracy than prior open-source baselines and achieves comparable visual quality at $36\times$ higher throughput for scalable world modeling.

📄 PDF Abstract BibTeX arXiv:2605.15178

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

2026-07-23 · Junsong Chen, Jincheng Yu, Yitong Li, Shuchen Xue 외 arxiv

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches…

Video Generation

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

2025-09-29 · Junsong Chen, Yuyang Zhao, Jincheng Yu, Ruihang Chu 외 arxiv

We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720x1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong…

Video GenerationVideo Alignment

An introductory guide to aligning networks using SANA, the Simulated Annealing Network Aligner

2019-11-22 · Wayne B. Hayes

Sequence alignment has had an enormous impact on our understanding of biology, evolution, and disease. The alignment of biological {\em networks} holds similar promise. Biological networks generally model interactions be…

CPU

Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention

2026-05-22 · Bingtian Qiao, Yue Shi, Yingjie Zhou, Yong Guo 외 arxiv

Real-world image super-resolution aims to recover high-quality images from complex and unknown real-world degradations. However, existing generative Real-ISR methods largely inherit the dense latent representations and q…

Image Super-Resolution

SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer

2026-05-28 · Yuyang Zhao, Yicheng Pan, Qiyuan He, Jincheng Yu 외 arxiv

Real-time streaming video-to-video editing (V2V) is critical for interactive applications such as live broadcasting and gaming, yet it remains a formidable challenge due to the stringent requirements for temporal consist…