paper-with-me

홈 › Papers

SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer

2026-01-23 · Tongcheng Fang, Hanling Zhang, Ruiqi Xie, Zhuo Han, Xin Tao, Tianchen Zhao, Pengfei Wan, Wenbo Ding, Wanli Ouyang, Xuefei Ning, Yu Wang arxiv

Diffusion Transformers have demonstrated remarkable performance in video generation. However, their long input sequences incur substantial latency due to the quadratic complexity of full attention. Various sparse attention mechanisms have been proposed. Training-free approaches are limited to moderate sparsity and thus yield only modest acceleration, whereas training-based methods can reach much higher sparsity but demand substantial data and computation. In this work, we propose SALAD, introducing a lightweight linear attention branch in parallel with the sparse attention. Leveraging a Multi-level Static-Dynamic Scaling Strategy to balance the two branches, our method attains up to 90% sparsity and 1.52-2.03x inference speedup across different models and sequence lengths, while maintaining generation quality comparable to the full attention baseline. Moreover, our finetuning process is highly efficient, requiring only 2,000 video samples, fewer than 1,600 training steps, and no more than 30 GPU hours with a batch size of 8.

📄 PDF Abstract BibTeX arXiv:2601.16515

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowingly

2025-11-01 · Wenya Xie, Shaochen, Zhong, Hoang Anh Duy Le 외 arxiv

Large Reasoning Models (LRMs) are often bottlenecked by the high cost of output tokens. We show that a significant portion of these tokens are useless self-repetitions - what we call "word salad" - that exhaust the decod…

SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing

2025-03-18 · CVPR 2025 1 · Seokhyeon Hong, Chaelin Kim, Serin Yoon, Junghyun Nam 외

Text-driven motion generation has advanced significantly with the rise of denoising diffusion models. However, previous methods often oversimplify representations for the skeletal joints, temporal frames, and textual wor…

DenoisingMotion Generation

SALAD: Source-free Active Label-Agnostic Domain Adaptation for Classification, Segmentation and Detection

2022-05-24 · Divya Kothandaraman, Sumit Shekhar, Abhilasha Sancheti, Manoj Ghuhan 외

We present a novel method, SALAD, for the challenging vision task of adapting a pre-trained "source" domain network to a "target" domain, with a small budget for annotation in the "target" domain and a shift in the label…

Active LearningDomain AdaptationImage SegmentationSemantic Segmentation

Sparse Attention with Linear Units

2021-04-14 · EMNLP 2021 11 · Biao Zhang, Ivan Titov, Rico Sennrich

Recently, it has been argued that encoder-decoder models can be made more interpretable by replacing the softmax function in the attention with its sparse variants. In this work, we introduce a novel, simple method for a…

DecoderDiversityMachine TranslationTranslation+1

RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers

2026-05-20 · Yuxi Liu, Zekun Zhang, Yixiang Cai, Renjia Deng 외 arxiv

Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, yet their $\mathcal{O}(L^2)$ attention complexity poses a formidable bottleneck for long-sequence synthesis. While recent sparse-linear at…

Video Generation