paper-with-me

홈 › Papers

Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention

2026-05-19 · Wenhu Zhang, Yiming Wu, Huanyu Wang, Yaoyang Liu, Huanzhang Dou, Senqiao Yang, Sitong Wu, Hanbin Zhao, Jiaya Jia arxiv

Diffusion Language Models (DLMs) enable globally coherent, bidirectional, and controllable text generation, offering advantages over traditional autoregressive LLMs, while scaling to ultra-long sequences remains costly. Many existing block-sparse attention methods select blocks by fixed sampling patterns over the high-resolution attention space, such as tail regions or anti-diagonal stripes. Such prior-driven sampling can miss salient tokens and introduce instability under distribution shifts. In this paper, we propose the Block Approximate Sparse Attention framework (BA-Att) with block-wise pre-downsampled operation, which identifies informative regions within a compact downsampled space, avoiding reliance on brittle positional priors. To analyze its theoretical behavior, we define an oracle post-downsample attention map and formalize the approximation error between pre- and post-downsample schemes. Based on this insight, we introduce a lightweight norm-sorting module and a covariance-compensated correction that approximates full covariance using diagonal QK variances, reducing computational complexity. Extensive experiments show that our operator achieves up to 6.95x acceleration over FlashAttention in attention computation, and maintains near full-attention performance at 50% sparsity across language models, multimodal language models, and video generation models, demonstrating strong efficiency and generalization.

📄 PDF Abstract BibTeX arXiv:2605.19726

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationText Generation

Similar Papers 제목 키워드 기반

From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs

2025-12-07 · Yuchuan Tian, Yuchen Liang, Shuo Zhang, Yingte Shu 외 arxiv

Diffusion Language Models (DLMs) enable fast generation, yet training large DLMs from scratch is costly. As a practical shortcut, adapting off-the-shelf Auto-Regressive (AR) model weights into a DLM could quickly equip t…

FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion

2026-02-05 · Zhuokun Chen, Jianfei Cai, Bohan Zhuang arxiv

Generating long-form content, such as minute-long videos and extended texts, is increasingly important for modern generative models. Block diffusion improves inference efficiency via KV caching and block-wise causal infe…

Causal InferenceVideo Generation

Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

2025-03-12 · Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang 외

Diffusion language models offer unique benefits over autoregressive models due to their potential for parallelized generation and controllability, yet they lag in likelihood modeling and are limited to fixed-length gener…

DenoisingLanguage ModelingLanguage Modelling

MAGE: All-[MASK] Block Already Knows Where to Look in Block Diffusion LLM

2026-02-15 · Omin Kwon, Yeonjae Kim, Doyeon Kim, Minseo Kim 외 arxiv

Block diffusion LLMs are an emerging paradigm for parallel language generation, but their KV caching makes memory access the dominant bottleneck in long-context inference. Sparse attention, which attends only to a small …

Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context

2026-06-25 · Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary 외 arxiv

Diffusion language models offer a promising alternative to autoregressive models due to their potential for parallel and iterative generation. However, existing approaches use a single network for both context representa…