paper-with-me

홈 › Papers

Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies

2025-10-07 · Chunsan Hong, Seonho An, Min-Soo Kim, Jong Chul Ye arxiv

Masked diffusion models (MDMs) have recently emerged as a novel framework for language modeling. MDMs generate sentences by iteratively denoising masked sequences, filling in [MASK] tokens step by step. Although MDMs support any-order sampling, performance is highly sensitive to the choice of which position to unmask next. Prior work typically relies on rule-based schedules (e.g., max-confidence, max-margin), which provide ad hoc improvements. In contrast, we replace these heuristics with a learned scheduler. Specifically, we cast denoising as a KL-regularized Markov decision process (MDP) with an explicit reference policy and optimize a regularized objective that admits policy improvement and convergence guarantees under standard assumptions. We prove that the optimized policy under this framework generates samples that more closely match the data distribution than heuristic schedules. Empirically, across four benchmarks, our learned policy consistently outperforms max-confidence: for example, on SUDOKU, where unmasking order is critical, it yields a 20.1% gain over random and a 11.2% gain over max-confidence. Code is available at https://github.com/chunsanHong/UPO.

📄 PDF Abstract BibTeX arXiv:2510.05725

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A2D2: Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding

2026-06-11 · Sophia Tang, Yuchen Zhu, Molei Tao, Pranam Chatterjee arxiv

Discrete diffusion models offer a simple and stable likelihood-based framework for sequence generation, recently extended to any-length settings via token insertion. Principled reward-guided fine-tuning for any-length di…

Adaptive Order Policies for Masked Diffusion

2026-05-29 · Jama Hussein Mohamud, Mohsin Hasan, Mirco Ravanelli, Yoshua Bengio arxiv

Masked diffusion models have seen great success in capturing data distributions over discrete sequences in domains such as text and proteins. These models generate data by iteratively unmasking tokens starting from a ful…

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors

2026-04-27 · Pengcheng Wang, Kaiwen Hong, Chensheng Peng, Katherine Driggs-Campbell 외 arxiv

Unlike chatbots, physical AI must act while the world keeps evolving. Therefore, the inter-chunk pause of synchronous executors are fatal for dynamic tasks regardless of how fast the inference is. Asynchronous execution …

Learning Unmasking Policies for Diffusion Language Models

2025-12-09 · Metod Jazbec, Theo X. Olausson, Louis Béthune, Pierre Ablin 외 arxiv

Diffusion (Large) Language Models (dLLMs) now match the downstream performance of their autoregressive counterparts on many tasks, while holding the promise of being more efficient during inference. One critical design a…

Reinforcement Learning

dgMARK: Decoding-Guided Watermarking for Diffusion Language Models

2026-01-30 · Pyo Min Hong, Albert No arxiv

We propose dgMARK, a decoding-guided watermarking method for discrete diffusion language models (dLLMs). Unlike autoregressive models, dLLMs can generate tokens in arbitrary order. While an ideal conditional predictor wo…