paper-with-me

홈 › Papers

Parallelism and Generation Order in Masked Diffusion Language Models: Limits Today, Potential Tomorrow

2026-01-22 · Yangyang Zhong, Yanmei Gu, Zhengqing Zang, Xiaomeng Li, Yuqi Ding, Xibei Jia, Yuting Shen, Zhenzhong Lan, Liwang Zhu, Weiping Liu, Junlin Zhou, Haisheng Liu, Zhong Xin Yu, Pengxin Luo, Donglian Qi, Yunfeng Yan, Junbo Zhao arxiv

Masked Diffusion Language Models (MDLMs) promise parallel token generation and arbitrary-order decoding, yet it remains unclear to what extent current models truly realize these capabilities. We characterize MDLM behavior along two dimensions -- parallelism strength and generation order -- using Average Finalization Parallelism (AFP) and Kendall's tau. We evaluate eight mainstream MDLMs (up to 100B parameters) on 58 benchmarks spanning knowledge, reasoning, and programming. The results show that MDLMs still lag behind comparably sized autoregressive models, mainly because parallel probabilistic modeling weakens inter-token dependencies. Meanwhile, MDLMs exhibit adaptive decoding behavior: their parallelism and generation order vary significantly with the task domain, the stage of reasoning, and whether the output is correct. On tasks that require "backward information" (e.g., Sudoku), MDLMs adopt a solution order that tends to fill easier Sudoku blanks first, highlighting their advantages. Finally, we provide theoretical motivation and design insights supporting a Generate-then-Edit paradigm, which mitigates dependency loss while retaining the efficiency of parallel decoding.

📄 PDF Abstract BibTeX arXiv:2601.15593

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference

2025-12-28 · Aiwei Liu, Minghua He, Shaoxun Zeng, Sijun Zhang 외 arxiv

Autoregressive (AR) generation is the standard decoding paradigm for Large Language Models (LLMs), but its token-by-token nature limits parallelism at inference time. Diffusion Language Models (DLLMs) offer parallel deco…

Unifying Masked Diffusion Models with Various Generation Orders and Beyond

2026-02-02 · Chunsan Hong, Sanghyun Lee, Jong Chul Ye arxiv

Masked diffusion models (MDMs) are a potential alternative to autoregressive models (ARMs) for language generation, but generation quality depends critically on the generation order. Prior work either hard-codes an order…

TRIMS: Trajectory-Ranked Instruction Masked Supervision for Diffusion Language Models

2026-04-01 · Lingjie Chen, Ruizhong Qiu, Yuyu Fan, Yanjun Zhao 외 arxiv

Diffusion language models (DLMs) offer a promising path toward low-latency generation through parallel decoding, but their practical efficiency depends heavily on the decoding trajectory. In practice, this advantage ofte…

Divide and Conquer: Accelerating Diffusion-Based Large Language Models via Adaptive Parallel Decoding

2026-02-27 · Xiangzhong Luo, Yilin An, Zhicheng Yu, Weichen Liu 외 arxiv

Diffusion-based large language models (dLLMs) have shown promising performance across various reasoning tasks, establishing themselves as an alternative to autoregressive large language models (LLMs). Unlike autoregressi…

On Powerful Ways to Generate: Autoregression, Diffusion, and Beyond

2025-10-07 · Chenxiao Yang, Cai Zhou, David Wipf, Zhiyuan Li arxiv

Diffusion language models have recently emerged as a competitive alternative to autoregressive language models. Beyond next-token generation, they are more efficient and flexible by enabling parallel and any-order token …