paper-with-me

홈 › Papers

Bidirectional Sparse Attention for Faster Video Diffusion Training

2025-09-01 · Chenlu Zhan, Wen Li, Chuyu Shen, Jun Zhang, Suhui Wu, Hao Zhang arxiv

Video diffusion Transformer (DiT) models excel in generative quality but hit major computational bottlenecks when producing high-resolution, long-duration videos. The quadratic complexity of full attention leads to prohibitively high training and inference costs. Full attention inefficiency stems from two key challenges: excessive computation due to the inherent sparsity of Queries and Key-Value pairs, and redundant computation as fixed sparse patterns fail to leverage DiT's dynamic attention. To overcome this limitation, we propose a Bidirectional Sparse Attention (BSA) framework for faster video DiT training, the first to dynamically sparsify both Queries and Key-Value pairs within 3D full attention, thereby substantially improving training and inference efficiency. BSA addresses these issues through two key components. Query sparsity is optimized by selecting the most informative query tokens via semantic similarity and with a dynamic spatial-time training strategy, while KV sparsity is achieved by computing a statistical dynamic threshold to retain only the most salient KV blocks for computation. Extensive experiments demonstrate that BSA significantly accelerates DiT training across long sequences, reducing FLOPs by up to 20x and achieving 17.79x faster attention training, while preserving or even surpassing the generative quality of full attention.

📄 PDF Abstract BibTeX arXiv:2509.01085

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Similarity

Similar Papers 제목 키워드 기반

Hierarchical Denoising For Multi-Step Visual Reasoning

2026-07-16 · Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou 외 arxiv

Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusio…

Visual ReasoningVideo Generation

MonarchRT: Efficient Attention for Real-Time Video Generation

2026-02-12 · Krish Agarwal, Zhuoming Chen, Cheng Luo, Yongqi Chen 외 arxiv

Real-time video generation with Diffusion Transformers is bottlenecked by the quadratic cost of 3D self-attention, especially in real-time regimes that are both few-step and autoregressive, where errors compound across t…

Computational EfficiencyVideo Generation

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

2026-06-09 · Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee, Joungbin Lee 외 arxiv

Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many denoising steps make them impractical for real-time inference. We pr…

Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-Clustering

2026-03-19 · Jiayi Luo, Jiayu Chen, Jiankun Wang, Cong Wang 외 arxiv

Diffusion Transformers (DiTs) achieve strong video generation quality but suffer from high inference cost due to dense 3D attention, motivating sparse attention techniques for improving efficiency. However, existing trai…

Video Generation

MAGE: All-[MASK] Block Already Knows Where to Look in Block Diffusion LLM

2026-02-15 · Omin Kwon, Yeonjae Kim, Doyeon Kim, Minseo Kim 외 arxiv

Block diffusion LLMs are an emerging paradigm for parallel language generation, but their KV caching makes memory access the dominant bottleneck in long-context inference. Sparse attention, which attends only to a small …